#1730. AmE-BrE 2006 Frequency Comparer ([2014-01-21-1]), #1739. AmE-BrE Diachronic Frequency Comparer ([2014-01-30-1]) ǡthe Brown family of corpora [2010-06-29-1]ε#428. The Brown family of corpora ѾաפȡˤѤѼ֤뤤̻ŪӥġäBrown family ȤС褦߷פԤޤ줿 ICE (International Corpus of English) ۵[2010-09-26-1]ε#517. ICE 7ϰѼ拾ѥפȡˡ1990ǯʹߤνդäդǼ줿100쵬ϤΥѥǡߤӲǽȤʤ褦˺Ƥ롥
ǡ긵ˤ ICE ΤCanada, Jamaica, India, Singapore, the Philippines, Hong Kong αѸѼ拾ѥ6оݤˡƱ褦ɽꡤǡ١ӤǽȤʤġȤˤĤƤϡ#1730. AmE-BrE 2006 Frequency Comparer ([2014-01-21-1]) Ȥ줿
#1730. AmE-BrE 2006 Frequency Comparer ([2014-01-21-1]) ǡ2006ǯνեƥȤԻƳѼ拾ѥҲ𤷡˴ŤӥġΥġʤ鵤ŤΤƱˡԻ졤ϤƱ100٤ the Brown family of corpora ʡ#428. The Brown family of corpora Ѿա ([2010-06-29-1])ˤϢȤСľ50ǯ֤ۤɤ̻ŪʱƴӤưפ˲ǽȤʤ롥
ǡεǾҲ𤷤 Professor Paul Baker - Linguistics and English Language at Lancaster University ˤ AmE06 BrE06 ˲äơեꥫѸɽ Brown (1961), Frown (1992)եꥹѸɽ LOB (1961), FLOB (1991) ɽФ碌ƥǡ١ѤλϡAmE-BrE 2006 Frequency Comparer ȤۤƱʤΤǡμ ([2014-01-21-1]) Ȥ줿ϤɽǤϡθиƥȤοٽ̤ϾʤƤꡤ100٤ɽˤȤɤƤΤǡAmE06 BE06 ˤĤԤξɬפʾˤϡAmE-BrE 2006 Frequency Comparer ɤ
Professor Paul Baker - Linguistics and English Language at Lancaster University ȤڡƤäBaker ԻѸ졦Ƹ쥳ѥ BE06 AmE06 ξȡФñꥹȤ롥ΥѥΤϡ桼ID᤹Сؤ CQP (Corpus Query Processor) system ꥢǤ롥
BE06 AmE06 ϡ2006ǯ˽Ǥ줿ꥹѼȥꥫѼνնѹեѥǤ롥Ի乽ϡ#428. The Brown family of corpora Ѿա ([2010-06-29-1]) ǾҲ𤷤 The Brown family ˽सƤꡤ500ƥȡ2000η100ۤɤεϤ
ơΥڡɤǤ BE06 Wordlist in WordSmith 5 format AmE06 Wordlist in WordSmith 5 format ʸФǤϤʤ˸ˤɽФ줾ǡ١ơѼθ٤ӤƤ AmE-BrE Frequency 2006 Comparer ʤġƤߤ
acknowledgment, acknowledgement, aging, ageing, aluminum, aluminium, analyze, analyse, apologize, apologise, armor, armour, behavior, behaviour, center, centre, civilization, civilisation, color, colour, defense, defence, disk, disc, endeavor, endeavour, favor, favour, favorite, favourite, fiber, fibre, flavor, flavour, fulfill, fulfil, gray, grey, harbor, harbour, honor, honour, humor, humour, inquiry, enquiry, judgment, judgement, labor, labour, license, licence, liter, litre, marvelous, marvellous, mold, mould, mom, mum, neighbor, neighbour, neighborhood, neighbourhood, odor, odour, organize, organise, pajamas, pyjamas, parlor, parlour, program, programme, realize, realise, recognize, recognise, skeptic, sceptic, specter, spectre, sulfur, sulphur, theater, theatre, traveler, traveller, tumor, tumour
ޤǤϡäֻ˴ؤƺΥѥˤӤϡ#708. Frequency Sorter CGI ([2011-04-05-1]) ѤꡤBNC Frequency Extractor ([2012-12-08-1]) ȡ#1322. ANC Frequency Extractor ([2012-12-09-1]) Ȥ߹碌ꡤthe Brown Family corpora ʻѤʤɡѼ拾ѥθӤˤн褷ƤΥġˤ¿ʴĶǤ
佼ˡ (suppletion) Ϲؿ⤿Ǥ롥go -- went -- gone, be -- is -- am -- are -- was -- were -- been, good -- better -- best, bad -- worse -- worst, first -- second -- third ʤɡʤƱηϤΤʤ˰ۤʤ촴ΤԻĤǤ롥ˤԵ§ζˤߤΤ褦˻פ뤫顤Ȥ櫓ؽԤܤˤȤޤ䤹
θؤˤƤϡ佼ˡؤδؿɬ⤯ʤ佼ˡ겼Ƹ椹뤳Ȥˤϸ³ȴƤ뤫ͳȤƤϡ(1) ñȯǤ뤳ȡ(2) ŪԵ§ʬԲǽǤ뤳ȡ(3) Ūʰ (paradigmatic pressure) ΩƤꡤŪ (analogy) Ϳʤȡʤɤ롥Ĥޤꡤġ佼ϡʸˡΤʤηŪ˰ȤǤùܤȤƸ̤ϿƤˤʤΤƤ롥̤ˤ줬ʤη֤äƤΤŪ (arbitrary) ǤΤƱͤˡ佼ʤη֤ʤΤŪǤꡤ꿼겼ǤϤʤȤȤ佼ˡħ֤ȤС찮ζˤƹ٤θˤʤȤȤ餤Ǥ롥
Hogg ϡ츫̷⤹褦˻פ "Regular Suppletion" Ȥ̾Ǥơ佼ˡθ³ǤˤꡤǽȤ佼ˡϡθ˽פʰ̣ĤȤHogg ϡѸˤ佼ˡˤꡤ4Ƥ롥
1ܤϡ"the replacement of one suppletion by another" 㤬ߤ뤳ȤǤ롥Ǥ˸űѸǤ yfel -- wyrsa -- wyrsta 佼ˡӤԤʤƤѸǤϸθ촴ؤꡤѸ bad -- worse -- worst ؤȻäߤǤϡԤ evil -- more evil -- most evil ȤʤäƤ롥yfel ϶ˤưŪʸְפŪʸžƤäȤˤꡤworse -- worst б븶ϰ̤夫Ūʸ bad ʤäȤȤˤʤ롥Hogg ϡűѸ *bæd ϥ֡äʸڤƤʤǤꡤºݤˤ14--18ʸڤ badder -- baddest ȤȤˡ§ŪѲƤϤȿ¬Ƥ (72) ޤDzǤϤ뤬evil bad ˤĤơӵѲϰʲΤ褦ŪѲФȤƤ (72)
evil worse worse, more evil more evil bad badder worse, badder worse
2ܤϡ"the preference for suppletion over regularity" Ǥꡤgo -- went ߤ뤳ȤǤ롥"to go" 佼ȤƸűѸ ēode Ѹ went ֤줿ȤϤ褯ΤƤ롥ˤä went θ߷ wend wended Ȥ§ŪʲȤѸ˾ˤʤäƤ롥went ǽפʤΤϡ§佼ޤȤ佼ˡηΤǤϤʤȤȤ䥹åȥǤϡӡ§ gaid gaed ߽Ф줿Ȥ¤⤢롥
3ܤϡ"the addition of regularity without disturbance of the suppletion" Ǥ롥űѸ bēon 3;ʣ߷1 syndon ϡĸ *-es ŪȯŸǤ synd synt Ȥ佼ˡư (preterite-present_verb) θʣ -on äΤǤ롥Ū˷ŪĤʤϤθ촴ˡŪˤղạ̈Ǥ롥ϡǿ줿 (2), (3) ȿ롥
4ܤϡ"the creation of a new regular inflection on the basis of suppletion" Ǥ롥űѸ bēon 1;θñ1 (e)am ϡAnglia Ǥϸ -m ˤŪ bīom ߽ФƱǤϤ줬ư˵ڤӡŪ1;θñ flēom (I flee) sēom (I see) ߽ФȤˤʤä衤佼ˤäʬϤʤϤ -m ޤDzȤˤʤ롥
Hogg (80--81) ϡʾΤ褦佼ˡθŪŪ餫ˤǡ"[S]uppletion is not merely a linguistic freak which does no more than give a small amount of pleasure to a rather giggling schoolboy. . . . [S]uppletion is a dynamic process." ȽҤ١佼ˡβǽõʤ顤ʸĤƤ롥
Hogg, Richard. "Regular Suppletion." Motives for Language Change. Ed. Raymond Hickey. Cambridge: CUP, 2003. 71--81.
#1424. CELEX2 ([2013-03-21-1]) ǾҲ𤷤ǡ١DzƤߤ褦ȹͤVersion 2 ǿ˲ä줿 (English Frequency, Syllables) Υ֥ǡ١ˤꡤѸǺǤ¿פΥ
ϡCELEX2 ΤȤˤʤäƤ륳ѥΤΤ7.26%130äե֥ѥФ줿٤Ǥꡤ٤ǤϤʤȡ٤ˤΤǤ롥Ĥޤꡤäդˤ뤢ñ٤⤱Сʬñ˴ޤޤ벻פ٤⤯ʤȤȤǤ롥㤨Сof "Ov" (= /ɒv/) ɽ벻ϡ4̤٤Ǥ롥ʤ̵ͭϹθ٤Ƥ롥
ʲΥꥹȤ˵벻ɽϡIPA ǤϤʤ CELEX ͤäɽʤΤǡбɽƤ

Ǥϡʲ˥ɽǥȥå50̤ޤǤǺܤ롥٤ñβפΤޤ̤ȿǤƤơޤꤪ⤷ɽǤϤʤΩĤȤ⤢뤫⤷ʤ
| Rank | Syllable | Frequency |
|---|---|---|
| 1 | eI | 72971 |
| 2 | Di: | 60967 |
| 3 | tu: | 31446 |
| 4 | Ov | 30108 |
| 5 | In | 29906 |
| 6 | &nd | 28709 |
| 7 | aI | 23822 |
| 8 | lI | 19728 |
| 9 | @ | 19566 |
| 10 | rI | 14356 |
| 11 | ju: | 12598 |
| 12 | dI | 12465 |
| 13 | D&t | 12118 |
| 14 | It | 11504 |
| 15 | wOz | 10834 |
| 16 | fO:r* | 9778 |
| 17 | Iz | 9517 |
| 18 | tI | 9161 |
| 19 | fO | 9042 |
| 20 | Sn, | 8969 |
| 21 | hi: | 8928 |
| 22 | r@n | 8638 |
| 23 | bi: | 8505 |
| 24 | bI | 7936 |
| 25 | nI | 7068 |
| 26 | wID | 7046 |
| 27 | On | 7030 |
| 28 | &z | 6919 |
| 29 | O:l | 6569 |
| 30 | h&d | 6240 |
| 31 | E | 6165 |
| 32 | bl, | 6021 |
| 33 | sI | 5836 |
| 34 | @U | 5824 |
| 35 | t@r* | 5687 |
| 36 | &t | 5652 |
| 37 | hIz | 5564 |
| 38 | bVt | 5416 |
| 39 | mI | 5397 |
| 40 | s@ | 5391 |
| 41 | nOt | 5357 |
| 42 | D@r* | 5339 |
| 43 | I | 5283 |
| 44 | tId | 5259 |
| 45 | DeI | 5162 |
| 46 | IN | 5063 |
| 47 | t@ | 5053 |
| 48 | s@U | 4974 |
| 49 | baI | 4894 |
| 50 | h&v | 4769 |
Please note that the English corpus used by CELEX for deriving these frequencies contains only 7.3% spoken material. This means there is a rather tenuous relationship between the full frequency figures, which are based on written forms, and the syllable frequencies, which merely refer to phonemic conversions of these graphemic transcriptions. Of course it could be argued that frequencies of syllables, as lexical sub-units, are less liable to get skewed from differences in medium than full words, but it has to be taken into account that NO FIRM EVIDENCE ABOUT SPOKEN FREQUENCIES can be derived from these data.
ñ٤˴ϢBetty Phillips ʤɡˤǡCELEX Ȥåǡ١ѤƤΤ뤳Ȥ롥꤫äƤ븦ǡ祳ѥ˴ŤǤפɬפˤʤäΤǡߤ350ɥ뤹뤳ιʥǡ١ꤷƤߤǤ2ǤǤꡤCELEX2 ȤƹǤ롥ʤʤͽۤƤʤäꤷ CD-ROM ˤϡLDC99T42 Ȥǡ١ޤޤƤˤ tagged Brown Corpus, Wall Street Journal, Switchboard tagged ʤ Treebank ϤΥѥäƤ롥
ơCELEX2 ˤϡѸä˴ؤʣΥǡ١ǼƤ롥줾Υǡ١ˤϡˡᡤ֡γƴ顤Ф (lemma) 뤤ϸ (wordform) ȤˡѥǤξǼƤ롥Ūˤϡ11Υǡ١ѲǽǤ롥
ect (English Corpus Types)
efl (English Frequency, Lemmas)
efs (English Frequency, Syllables)
efw (English Frequency, Wordforms)
eml (English Morphology, Lemmas)
emw (English Morphology, Wordforms)
eol (English Orthography, Lemmas)
eow (English Orthography, Wordforms)
epl (English Phonology, Lemmas)
epw (English Phonology, Wordforms)
esl (English Syntax, Lemmas)
Ф줢뤤ϸȤ token ٤μФ˶ǡ١ȤǧǹºݤˤϡޤޤƤμ϶äۤ˭٤ǡ11Υǡ١٤Ƥ碌եɿϤΤ250ʾ˵ڤ֡Կ efl 52,447ԡefw 160,595ԤȤ礵Ѥ SQLite DB 館顤̤ˤ90MBĶƤޤä
CELEX2 ΥϡˤĤƤ Oxford Advanced Learner's Dictionary (1974) ڤ Longman Dictionary of Contemporary English (1978) ǤꡤپˤĤƤ 1790줫ʤ COBUILD/Birmingham corpus Ǥ롥Υѥιϡ1660 (92.74%) եѥ130 (7.26%) äեѥǡԤ284ƥȤΤ44ƥ (15.49%) ꥫѸǤ롥ΥꥫѸϤۤȤɤꥹѸֻľƤ뤳Ȥդ
CELEX2 ˤ "lemma" ϡʲ5˰¸롥
(1) orthography of the wordforms: peek vs peak
(2) syntactic class: meet (adj.) vs meet (adv.)
(3) inflectional paradigm: water (v.) vs water (n.)
(4) morphological structure: rubber (someone or something that rubs) vs rubber (the elastic substance)
(5) pronunciation of the wordforms: recount [ˈriː-kaʊnt] vs recount [rɪ-ˈkaʊnt]
äơ̾ۤʤ lexeme Ȥư bank (ڼˤ bank ʶԡˤʤɤϡCELEX2 ǤƱ lemma ȤưƤΤդɬפǤ롥
Τ褦 CELEX2 ˶Ϥʸ٥ǡ١¾ˤٸ˻ǡ١ġ¸ߤ롥ܥ֥ǿ줿ΤȤƤϡfrequency statistics lexicology γƵ䡤ä˰ʲεͤˤʤ
#308. Ѹκѱñꥹȡ ([2010-03-01-1])
#607. Google Books Ngram Viewer ([2010-12-25-1])
#708. Frequency Sorter CGI ([2011-04-05-1])
#1159. MRC Psycholinguistic Database Search ([2012-06-29-1])
Baayen R. H., R. Piepenbrock and L. Gulikers. CELEX2. CD-ROM. Philadelphia: Linguistic Data Consortium, 1996.
[2012-02-13-1]ε#1022. ѸγƲǤ١פdzǧǤ褦ˡ/ʊ/ ϱѸû첻ǤΤʤǺǤ٤㤤ΤǤ롥ֻȤƤϡŵŪ <oo> <u> ɽ蘆벻ǤԤ /uː/Ԥ /ʌ/ ȤƼ¸뤳Ȥ¿ΤǡºݤˤϤۤɸʤŪˤϤɤΤ餤Τ
#1191. Pronunciation Search ([2012-07-31-1]) ǡ/ʊ/ + ҲפǽñäƤߤ "^[^AEIOU]*(?<!Y )UH[012]? [^AEIOUHWY]+$" Ƥߤȼ159줬ä
bloor, book, book's, booked, books, books', boor, boord, boors, bourque, brook, brook's, brooke, brooke's, brookes, brooks, brooks's, bruehl, bull, bull's, bulls, bulls', cook, cook's, cooke, cooked, cooks, could, crook, crooke, crooks, duerr, duerst, flook, fluhr, fooks, foor, foot, foote, foote's, foots, fuhr, fuld, full, full's, fulp, fults, fultz, gloor, good, good's, goode, goods, gook, hood, hoods, hoofed, hoofs, hook, hook's, hooke, hooked, hooks, hooves, joong, jure, kook, kooks, koors, kuehl, kuhrt, look, looked, looks, loong, luehrs, luhr, luhrs, lure, lured, lures, mook, moor, moore, moore's, moored, moores, moors, muhr, nook, nooks, poor, poor's, poore, poors, pull, pulled, pulls, puls, pultz, put, puts, rook, rooke, rooks, routes, ruehl, ruhr, schnooks, schnoor, schoof, schook, schultz, schulz, schulze, schuur, shook, should, shultz, shure, snook, snooks, soot, soots, spoor, spoor's, stood, stroock, stuhr, suire, sure, took, tooke, tookes, tour, tour's, toured, tours, ture, uhr, wolf, wolf's, wolfe, wolfe's, wolff, wolves, wood, wood's, woods, wool, woolf, wools, woong, would, wuertz, wulf, wulff, zook, zuehlke
ǡбĹ첻 /uː/ 첻 /ʌ/ ñϡƱͤξ︡ "^[^AEIOU]*(?<!Y )UW[012]? [^AEIOUHWY]+$" "^[^AEIOU]*(?<!Y )AH[012]? [^AEIOUHWY]+$" ˤС줾596졤966줬ҥåȤꤵ줿ĶˤĴǤϤ뤬Ū /ʊ/ ñ줬ʤȤ狼롥
171Τ<oo> = /ʊ/ δطΤ95줢롥ֻȯδططʤˤϡ#547. <oo> ֻб3ȯ ([2010-10-26-1]) ӡ#1297. does, done 첻 ([2012-11-14-1]) ǽҤ٤̤ꡤĹ첻û첻ȤѲä
βѲϲΤΤǤϤ뤬;ȤϸǤ⾯θˤƻȯŪ˸롥㤨Сȯɤΰ ([2010-08-28-1]) dzǧǤ¤ꡤroom, bedroom, broom ȯϡĹ첻û첻δ֤ɤLPD Preference polls ˤСɤʬۤϰʲ̤ꡥ
| BrE /ruːm/ | BrE /rʊm/ | AmE /ruːm/ | AmE /rʊm/ | |
|---|---|---|---|---|
| room | 81% | 19 | 93 | 7 |
| bedroom | 63 | 37 | - | - |
| broom | 92 | 8 | - | - |
ѸβäǤϡղõ䤬褯Ȥ롥ŪˤɤΤ餤褯ȤΤ⤽Ū˵ʸϤɤΤ餤٤ΤΤʤǡղõϤɤ줯餤γΤΤ褦ʵ顤ޤ٤ Biber et al. LGSWE Ǥ롥
ǽˤĤƤϡp. 211 ˲ͿƤ롥οˤƤĴCONV(ERSATION) Ǥ401ĵ䤬ޤޤƤȤåѥǤϡž̾塤䤬ȿǤƤǽ⤯ºݤˤϿͰʾ٤ǵʸƤϤǤ롥ƥȥפǤС礭 FICT(ION) ³NEWS ACAD(EMIC) Ǥϵʸ٤ϸ¤ʤ㤤
ˡƥ֥ѥˤơʸΤˤղõϤɤΤ餤p. 212 ˷ǺܤƤ̤ʲΤ褦ˤޤȤĤ100%ȤʤɽǤ롥
| (* = 5%; ~ = less than 2.5%) | CONV | FICT | NEWS | ACAD | |
|---|---|---|---|---|---|
| independent clause | wh-question | **** | ******* | ********* | ********** |
| yes/no-question | ***** | ***** | ******* | ******* | |
| alternative question | ~ | ~ | ~ | ~ | |
| declarative question | ** | * | ~ | ~ | |
| fragments | wh-question | * | ** | ** | * |
| other | *** | *** | * | * | |
| tag | positive | * | ~ | ~ | ~ |
| negative | **** | * | ~ | ~ | |
Tag questions (i.e., regular questioning expressions tagged onto a sentence) exist in both American and British English, with British speakers perhaps using them more than Americans: "That's not very nice, is it?" Peremptory and aggressive tags tend to be used more in British English than in American English: "Well, I don't know, do I?" (192)
ǰʤ顤Biber et al. Ǥղõ٤αƺΤ뤳ȤϤǤʤӡƤΥѥĴ٤ɬפ
Schmitt, Norbert, and Richard Marsden. Why Is English Like That? Ann Arbor, Mich.: U of Michigan P, 2006.
Biber, Douglas, Stig Johansson, Geoffrey Leech, Susan Conrad, and Edward Finegan. Longman Grammar of Spoken and Written English. Harlow: Pearson Education, 1999.
Cheshire (115) ɤǤơѸ˴ؤ뵭ҤȤơäˤ¿ȤȤڤľŪˤϳΤˤΤ褦˻פ뤬ҴŪդϤΤȡLGSWE äƤߤȡϢ뵭Ҥ pp. 159--60 ˸Ĥä
ˤ͡ʼब뤬4ĤλѰΤ줾ˤĤơѥѤ "Distribution of not/n't v. other negative forms" Ĵ̤Ƥ100դɽǼ

| not/n't | other negative forms | |
|---|---|---|
| CONV | 19500 | 2500 |
| FICT | 9500 | 4000 |
| NEWS | 4500 | 2000 |
| ACAD | 3500 | 1500 |
ε#1321. BNC Frequency Extractor ([2012-12-08-1]) ˰³ANC (American National Corpus) ˴ŤɽANC Second Release Frequency Data Υڡ˸ƤΤǡ"ANC Frequency Extractor"
# եƥȤǡƺȤ "diarrhoea" vs. "diarrhea" ֻ٤ǧ
select * from written where word like "diarrh%"
# եƥȤǡƺȤ "judgement" vs. "judgment" ֻ٤ǧʤ¾[2009-12-27-1]ε#244. ֻαƺΥꥹȡפֻǤ椯Ȥ⤷
select * from written where word like "judg%ment%"
# -ly ǽʤõflat adverb ⤷ʤõ
select * from anc where lemma not like "%ly" and pos like "RB%"
# -s ǽõadverbial genitive ̾Ĥ⤷ʤõ
select * from anc where pos like "RB%" and word like "%s"
# ñ̾ʣ̾ token Ӥ written subcorpus spoken subcorpus ǡ[2011-06-07-1]ε#771. ̾ñʣ١פȡ
select pos, sum(freq) from written where pos in ("NN", "NNS") group by pos
select pos, sum(freq) from spoken where pos in ("NN", "NNS") group by pos
select pos, sum(freq) from anc where pos in ("NN", "NNS") group by pos
ANC ͭȴ褵줿 OANC (Open American National Corpus) ̵ANC ڤ OANC ˤĤƤϡ#708. Frequency Sorter CGI ([2011-04-05-1]) #509. Dracula ˸ whilst (2) ([2010-09-18-1]) ȡ
"BNC Frequency Extractor" "ANC Frequency Extractor" Ȥ߹碌ƻȤСäαƺˤĤ٤δñĴǤ롥
Adam Kilgarriff Ƥ BNC database and word frequency lists 顤Ф첽Ƥʤɽ (unlemmatised lists) ɤǤ褦˥ǡ١館
# եƥȤǡƺȤ "diarrhoea" vs. "diarrhea" ֻ٤ǧ
select * from written where word like "diarrh%"
# s ǻϤޤʬι⤤
select * from variances where word like "s%" order by variance desc limit 100
# 첻Ѱۤʣñ١cf. #708. Frequency Sorter CGI([2011-04-05-1]) Ǥ lemma ä
select * from bnc where word in ("foot", "goose", "louse", "man", "mouse", "tooth", "woman") and pos = "nn1" order by freq desc
# 첻Ѱۤʣ
select * from bnc where word in ("feet", "geese", "lice", "men", "mice", "teeth", "women") and pos = "nn2"
# POSǤޤȤ٤ι⤤ˡä 'demog'
select pos, sum(freq) from demog group by pos order by sum(freq) desc
# Ǥ¿Ȥ̾
select * from variances where pos like "n%" order by variance desc limit 100
# Ǥ¿Ȥƻ
select * from variances where pos like "aj%" order by variance desc limit 100
ʤФ첽Ƥɽ (lemmatised list) ˤĤƤϡ٤ˤ800ʾ帽롤6318̤ޤǤθФΤߤ˸ꤵƤꡤθġϡ#708. Frequency Sorter CGI ([2011-04-05-1]) ȤƼƤ롥Ϣơ#956. COCA N-Gram Search ([2011-12-09-1]) ⻲ȡ
ε#1286. ֲѲΰۤʤ2ưŤ ([2012-11-03-1]) ǾҲ𤷤 Hooper ʸǤϡĴ1ĤȤưζܹԡʶѲưμѲˤ夲ƤHooper εñǤ롥ܹԤˤʿ (analogical leveling) ŵǤꡤ٤㤤ư줫˰ܹԤ뤲ƤΤȤ
Hooper ĴоݤȤưϸűѸζѲI, II, IIIͳ褹ưΤߤǤꡤθѸˤپˤĤƤ Kučera and Francis ɽȤƤ롥ٷ lemma ñ̤ǤֻΤߤȤФǤꡤdrive, ride ʤɤθʲɽ * դƤΡˤˤĤʻζ̤θƤʤӺʤΤޤǯʾˤ錄ѲˤƤȤˡѸˤ٤ΤߤȤƤ褤ΤȤ[2012-09-21-1]ε#1243. ٤θ̻ŪΤˡסˤˤĤƤڴŪǤ (99) ΤȤơ᤹Τ˻ͤޤǤˤȤâɬפʲ Hooper (100) ɽ䤹ѤΤǤ褦

ΤˤΤ褦˸ȡܹԤФưΤȤ٤Ūˤä㤤Ȥ狼롥Ϣơkeep, *leave, *sleep *creep, *leap, weep ˤĤơ3줬ŪʲݻƤΤФơ3ˤϼŪ creeped, leaped, weeped ΰ۷ǧȤԤ٤Ϥ줾 531, 792, 132 ФƸԤϤ줾 37, 42, 31 Ȥ (Hooper 100) ͤޤǤˤȤϤäƤ⡤ȤƤ餫Τ褦˻פ롥
ưζܹԤϱѸˤˤƴŪǤꡤܥ֥Ǥ#178. ưε§Ѳά ([2009-10-22-1]) #527. Ե§Ѳưε§®٤ٻɸ2ȿ㤹? ([2010-10-06-1]) #528. ˵§ư wed !? ([2010-10-07-1]) ʤɤǿƤƳȤ狼äƤʤȤ¿ξܺ٤ʸ椬ؤ롥
Hooper, Joan. "Word Frequency in Lexical Diffusion and the Source of Morphophonological Change." Current Progress in Historical Linguistics. Ed. William M. Christie Jr. Amsterdam: North-Holland, 1976. 95--105.
Ѳȸ٤ȤδطˤĤƤϡPhillips θҲ𤷤ʤ顤#1239. Frequency Actuation Hypothesis ([2012-09-17-1]) #1242. -ate ưζܹԡ ([2012-09-20-1]) Ǽ夲Ƥ#1265. ٤ȲѲνδط˵ŤƤ Schuchardt ([2012-10-13-1]) ǿ줿褦ˡ˲ŪѲϹٸ줫ϤޤȤȤ19ŦƤդ (analogy) δؤֲŪѲٸ줫ϤޤȤȤ⡤ۤƱ Herman Paul ˤäƻŦƤHooper 95)
Hooper ϡ1984ǯʸǡ٤Ȥ顤벻Ѳȹͤ븽Ѹˤ schwa-deletion memory ʤɤ2첻ˤˤʿȹͤưμѲĴ2Ω "phonetic change tends to affect frequent words first, while analogical leveling tends to affect infrequent words first" (101) ٻPhillips Ϥñ2ΩˤäƤǤʤΤ뤳ȤƤ뤬ΩνȯȤ뤳ȤϺǤ
Τ褦٤ȲѲδطˤäƤ褦˸ Hooper ¤ΤȤüԤϸپ˥ǤʤϤȹͤƤ (102)
I do not think the relative frequency of words is a part of native speaker competence, so I would not propose to make the rule sensitive to word frequency.
Ǥ٤ȲѲοʹԽشط뤳Ȥǧ³ΤǤСξԤϡüԤθǽϤǤϤʤ챿ѤΤʤˤȤȤˤʤΤHooper ϻҶθĤ褦ȤƤ褦
٤ȲѲν˴ؤΤŪʰյϡʷ֡˲Ѳˤϸΰۤʤʣμब뤳Ȥˤ롥ѲνۤʤȤȤϡ餯ѲΥᥫ˥बۤʤȤȤǤꡤѲθۤʤȤȤǤϤʤȤСѲν狼СѲưŤ狼뤳Ȥˤʤ롥ϤΤ褦ʸΰۤʤѲˡֽ벻ѲפŪʷֲѲפȤ٥ŽդƤϤ٤ʬबɬפHooper (103) ηƤ
. . . if it turned out that vowel shifts and some other phonetic changes affect infrequent forms before frequent forms, then we would have an interesting indication that phonetic changes arise from different sources, and furthermore, if my hypotheses are correct, a way of determining which types of changes are traceable to which source. Thus it appears that lexical diffusion, studied in terms of word frequency, may turn up some interesting evidence concerning the source of morpho-phonological change. (103)
Hooper, Joan. "Word Frequency in Lexical Diffusion and the Source of Morphophonological Change." Current Progress in Historical Linguistics. Ed. William M. Christie Jr. Amsterdam: North-Holland, 1976. 95--105.
Phillips, Betty S. "Word Frequency and the Actuation of Sound Change." Language 60 (1984): 320--42.
Phillips, Betty S. "Word Frequency and Lexical Diffusion in English Stress Shifts." Germanic Linguistics. Ed. Richard Hogg and Linda van Bergen. Amsterdam: John Benjamins, 1998. 223--32.
٤óȻ (lexical diffusion) οʹԽδطˤĤƤϡ[2012-09-17-1], [2012-09-20-1], [2012-09-21-1]γƵǰäƤ19ǯʸˡ (Neogrammarians) ˤСѲ "phonetically gradual and lexically abrupt" ǤȤȤQäǯθóȻθˤꡤѲˤѲŵŪ˸ "lexically gradual" βǧ褦ˤʤäƤƤ롥֤Ǥ첻Ǥ졤ѲΤʤˤϽȵڤƤ椯ΤȤθϡѲθȤƤ (analogy) Ȥ褯Ѳ֤Ƥ Neogrammarians ⤷ΤȤΤäξȤ
˿ʹԤ벻Ѳǧ硤줬ɤΤ褦ʽǿʹԤΤȤȤʤ롥ǡ٤ȤƤƤΤδθϤޤ˽ФǤ롥ȤȤȤǤСƤϰճʤƤǯʸˡɤй Hugo Ernst Maria Schuchardt (1842--1927) 1885ǯȤᤤʳǡ٤ȲѲνܤդƤΤʲϡPhillips (321) ˷ǺܤƤ Schuchardt ΰѡʱˤǤ롥
The greater or lesser frequency in the use of individual words that plays such a prominent role in analogical formation is also of great importance for their phonetic transformation, not within rather small differences, but within significant ones. Rarely used words drag behind; very frequently used ones hurry ahead. Exceptions to the sound laws are formed in both groups.
Neogrammarian λˤơ٤ȲѲνܤ Schuchardt ״˶äʤμϡθˤüξǤꡤƥΤȤ˭٤dzΤʴפΤäޥؤʬबȤꤸƤȤȴط롥ϡ¤λˤ Neogrammarian βˡ§ȿ㤬˭٤ˤ뤳ȤϤäǧƤꡤˡ§ˤˤκƷǤϸοʤȤ˴Ƥ²ȤͤˤŪǡ˺ȾȻβꤷƤäκؤδؿϲ̤ƤʤŪʸ椬ɾ褦ˤʤäΤϡ1塤쥪츦θ1960ǯΤȤǤ롥
Schuchardt Τ褦ʸۤäƤȤΤС嵭ΰѤ⼫ǤSchuchardt ȥ쥪ؤȯŸȤδطˤĤƤϡ (188--96) ܤ
Phillips, Betty S. "Word Frequency and the Actuation of Sound Change." Language 60 (1984): 320--42.
Schuchardt, Hugo. Über die Lautgesetze: Gegen die Junggrammatiker. 1885. Trans. in Shuchardt, the Neogrammarians, and the Transformational Theory of Phonological Change. Ed. Theo Vennemann and Terence Wilbur. Frankfurt: Altenäum, 1972. 39--72.
ɧظؤȤϲ١ȽŹҴȿӡ1993ǯ
ε#1242. -ate ưζܹԡ ([2012-09-20-1]) #1239. Frequency Actuation Hypothesis ([2012-09-17-1]) Ǽ夲 Phillips θΤ褦ˡ٤θѲθˤ¿ʴؿƤ뤬ˡѤʵȤơ٤켫Τ̻ŪѤȤ¤ɤΤ褦˹ͤФ褤ΤȤ꤬롥θǤϤʤäΤ뤤ʬͤˤϡ1--2λ礭ǤϤʤľ롥1λǤϤɤ2ǤϤɤȹͤȡɤޤľΤϤʤϤʤPhillips (225--26) ϡˤĤƼΤ褦˳ڴѤƤ롥
The words' frequencies are based on present-day English, but the general pattern of relative frequencies probably holds for the English in our data base (1755--1993) as well. For example, I would be very surprised if the 3-syllable verbs with CELEX frequencies over 100 --- concentrate, demonstrate, illustrate, contemplate, compensate, designate, and alternate --- were not also much more common in 1755 than those with frequencies of 0 --- altercate, auscultate, condensate, defalcate, eructate, exculpate, expuergate, extirpate, fecundate, etc.
2;λˤƤʤ顤٤100ʾθ0θ٤ȤΤ绨Ĥˤ褦˻פ롥ΤˡPhillips ϼºݤʬϤǤ101ʾ塤10--1001--10ȤӤʬѤƤꡤ绨Ĥپ绨ĤʤޤޤѤ뿵ŤϼƤ롥⤷ä10--100դ٥٥θܺ٤Ĵ٤褦ȤΤǤС2δ֤ˤʤ٤ѲƤǽϤ롥Phillips ʤ餺Ȥ⡤٤Ѥ̻Ū˴ؿï⤬ͤϤ
˻פĤñʲƤϡƻɽǤ礭ʥѥѤɽ뤳ȤǤ롥ƤȤƤñºݤ˿ԤΤϰ֤֤⤫롥ֻٸꤷѸǤСѥѰդɽμưǤѸǤֻ variation 椨 lemmatise Ƥʤ¤ϸФñ̤ǤɽҤޤ夬ŤʤФʤۤɡѥ˴ޤޤƥȤ representativeness Ͽˤʤ롥ӤäݤɽǤ⡤ʤϤۤ褤ƤߤȻפäƤ롥뤤ϡˤäƤϤǤˤ
ʤѤˤ CELEX Ȥñǡ١ϡѸθ֤˴ؤŪʸǤ褯ȤƤΤǤ롥ܺ٤ϡCELEX2 ȡޤ٤̻֤δطˤĤƤϡ[2012-05-03-1]ε#1102. Zipf's law ȸοաפȡ
Phillips, Betty S. "Word Frequency and Lexical Diffusion in English Stress Shifts." Germanic Linguistics. Ed. Richard Hogg and Linda van Bergen. Amsterdam: John Benjamins, 1998. 223--32.
[2012-09-17-1]ε#1239. Frequency Actuation Hypothesisפǡ -ate ưζưƤƤȤѲ˿줿̾ư (diatone Ʊͤ˸ߤʹζܹԤǤꡤóȻ (lexical_diffusion) ȤƤܤͤ롥
ΰܹԤˤĤƤ Phillips óȻδ鸦椷Ƥ뤬OED "contemplate, v" ˤޤȤޤä⤬ΤǡҲ𤷤褦Danielsson (271--72) Ǥ⡤OED Τβս꤬Ƥ롥
In a few rare cases (Shakes., Hudibras) stressed 'contemplate in 16--7th c.; also by Kenrick 1773, Webster 1828, among writers on pronunciation. Byron, Shelley, and Tennyson have both modes, but the orthoepists generally have con'template down to third quarter of 19th c.; since that time 'contemplate has more and more prevailed, and con'template begins to have a flavour of age. This is the common tendency with all verbs in -ate. Of these, the antepenult stress is historical in all words in which the penult represents a short Latin syllable, as ac'celerate, 'animate, 'fascinate, 'machinate, 'militate, or one prosodically short or long, as in 'celebrate, 'consecrate, 'emigrate; regularly also when the penult has a vowel long in Latin, as 'alienate, 'aspirate, con'catenate, 'denudate, e'laborate, 'indurate, 'personate, 'ruinate (L. aliēno, aspīro, etc.). But where the penult has two or three consonants giving positional length, the stress has historically been on the penult, and its shift to the antepenult is recent or still in progress, as in acervate, adumbrate, alternate, compensate, concentrate, condensate, confiscate, conquassate, constellate, demonstrate, decussate, desiccate, enervate, exacerbate, exculpate, illustrate, inculcate, objurgate, etc., all familiar with penult stress to middle-aged men. The influence of the noun of action in -ation is a factor in the change; thus the analogy of ,conse'cration, 'consecrate, etc., suggests ,demon'stration, 'demonstrate. But there being no remonstration in use, re'monstrate, supported by re'monstrance, keeps the earlier stress.
Ĥޤꡤ3ʾθˤƤϡŪˤ penult ι˱ƶ penult antepenult Ūˤϡpenult ˻ҲˤϡŪˤϤβ˶ȤḽѸˤơб̾ -ation ζѥˤȤŤ䤬Ưᤫζ˰ĺءantepenult ؤȰܹԤƤƤȤΤǤ롥
Danielsson OED ˤ3ʾθˤĤƤθڤʤPhillips 2ˤĤƤĴ̣Ȥˡ2 -ate ưʤ penult IJΤΡˤǤϡȿФζܹԤäƤȤfrustrate, dictate, prostrate, pulsate, stagnate, truncate ʤɤθǤϡŪˤ penult ˶Ѹˤ ultima ˶۷ƤƤʺǽ3ˤĤƤ夷ˡơζܹԤˤĤƤ⡤Phillips ٤ι⤤ΤѲƤƤȤ¤ͤߤ (226--28)
٤㤤ΤѲƤ Phillips μĥ̾ưȡΩ̤Ǥ롥٤ȸóȻοʹԽȤˡʰФꤸƤ롥
Phillips, Betty S. "Word Frequency and Lexical Diffusion in English Stress Shifts." Germanic Linguistics. Ed. Richard Hogg and Linda van Bergen. Amsterdam: John Benjamins, 1998. 223--32.
Danielsson, Bror. Studies on the Accentuation of Polysyllabic Latin, Greek, and Romance Loan-Words in English. Stockholm: Almqvist & Wiksell, 1948.
óȻ (lexical diffusion) ȤƿʹԤ벻Ѳƻڤ٤ؤƤ餷ȤϡŤ19ŦƤºݤˡPhillips (1984: 321) ˵Ƥ褦ˡ٤ι⤤줫Ѳ뤲ȤѲϿڤƤǡ٤㤤줫Ѳ뤲ǧƤꡤ٤ȸóȻνδطˤĤƤϡޤ˵䤬¿ˤĤơPhillips ϡꥫѸˤ glide deletion Ѹ unrounding Ѹ̾ưdiatonic stress shift; diatone εȡˤȤ٤㤤˿ʹԤȤ3ĤβѲ夲ơ"Frequency Actuation Hypothesis" ϡ"physiologically motivated sound changes affect the most frequent words first; other sound changes affect the least frequent words first" (1984: 336) ȤΤǤԤ surface phonetic form ƯѲԤ underlying phonetic form ƯѲؤˡ
Phillips 1998ǯ -ate ǽưζ֤ΰư˴ؤ븦ˤơŪưŤƤʤѲͽۤ褦٤㤤ˤϿʤޤष٤ι⤤˿ʤǤ뤳Ȥ餫ˤǡ Frequency Actuation Hypothesis
[F]or segmental changes, physiologically motivated sound changes affect the most frequent words first; other sound changes affect the least frequent words first. For suprasegmental changes, changes which require analysis (e.g., by part of speech or by morphemic element) affect the least frequent words first, whereas changes which eliminate or ignore grammatical information affect the most frequent words first. (1998: 231)
ĤޤꡤΰưΤ褦ĶʬβѲ˴ؤƤϡüԤˤʬϤ뤫ʤǡ٤ȽδطžȤ櫓Ǥ롥ʤʤΤˤĤơPhillips Bybee (117--19) "lexical strength" ȤͤФƤ롥
ɬ⤳εǼƤʤޤPhillips μĥȤϰۤʤꡤ̾ư夬٤ι⤤˿ʹԤȤǡȼƤ롥٤ѲνˤĤƤθϽ˽ФǤꡤ;Ϥ¿ʬ˻ĤƤ롥
Phillips, Betty S. "Word Frequency and the Actuation of Sound Change." Language 60 (1984): 320--42.
Phillips, Betty S. "Word Frequency and Lexical Diffusion in English Stress Shifts." Germanic Linguistics. Ed. Richard Hogg and Linda van Bergen. Amsterdam: John Benjamins, 1998. 223--32.
Bybee, Joan L. Morphology: A Study of the Relation between Meaning and Form. Amsterdam: John Benjamins, 1985.
#1108. 쵭̵ͭ ([2012-05-09-1]) #1110. Guiraud ˤؤι ([2012-05-11-1]) ǻȤ̣ؼԤ Guiraud ϡ (information_theory) θؤؤαѤˤؿηϤ쵭Τ;١ѤʤɤͻƤ롥
1954ǯʸɤߡ¿μŪƶ줿㤨С˥ե˥ե١ĹδطˤĤƼΤ褦˽Ҥ٤Ƥ (128) ǽϥ˥ե˥ե֡סкǤû˥եǤ٤ι⤤˥ե˳Ƥ롥줫顤˥եˡֶưסˡѹäס
Υ˥եȥ˥եߴطްդΤϡ餫ͳ٤̣֤ѲƤ椯ȡޤݤƤξԤδ֤ζѹդ뤿ˡηϤĴǽȯưѹդȤȤȤǤ롥̤θСѲϡãθΨ¤ݤ¤ˤƵȤȤˤʤ롥ηϤηϤ1ĤǤʾ塤˴ؤ̸ǤָΨפ˽虜ʤȤˤʤ
ǤϡָΨפ졤ְ̣פϼξݤΤ̤̣ȤȤƤ Guiraud ϡΤ褦ˡǾθ̣Ѥ˳褫ȹͤƤ롥ҴŪ˿ɽ蘆٤ĹȤɸѤơܤ˸ʤ˥եȥ˥եδطõΤǤϤʤ
La relation coût/information (ou forme/fréquence) traduit objectivement ces rapports entre le signe et le concept et permet de poser en termes objectifs le problème de la signification. (128)
ѡʤ뤤Ϸ֡١ˤδطϥ˥եȥ˥եδ֤ΤδطҴŪɽ魯ΤǤꡤ̣ѤҴŪ뤳ȤǽˤƤ롥
Guiraud, P. "Langage et communication. Le substrat informationnel de la sémantisation." Bulletin de la société de linguistique de Paris 50 (1954): 119--33.
ε[2012-06-28-1]ǾҲ𤷤Ѹåǡ١ MRC Psycholinguistic Database ܥ֥夫ʰġºݤˤϸġȤϡMRC Psycholinguistic Database ѤȡʤȤǤȤȤǥǤˤϷ̤10ԤΤߤ˸ꤷƤ롥ܳŪʻѤˤϡڡǡ١ȸץɤ뤫־Υե (Online search (answers limited to 5000 entries) or Online search (limited search capabilities)) ɤ
# ʸǸäʬ
select NLET, count(NLET) from mrc2 group by NLET;
# ǿǸäʬ
select NPHON, count(NPHON) from mrc2 group by NPHON;
# Ǹäʬ
select NSYL, count(NSYL) from mrc2 group by NSYL;
# -ed ǽƻٽ
select WORD, K_F_FREQ from mrc2 where WTYPE = 'J' and WORD like '%ed' order by K_F_FREQ desc;
# 2̾졤ƻ졤ưѥȤʬ (#814. ̾ưʤ̷ư ([2011-07-20-1]) ڤӡ#801. ̾ưε (3) ([2011-07-07-1]) ȡ
select WTYPE, STRESS, count(*) from mrc2 where NSYL = 2 and WTYPE in ('N', 'J', 'V') group by WTYPE, STRESS;
# <gh> ֻǽꡤ/f/ ȯǽ
select distinct WORD, DPHON from mrc2 where WORD like '%gh' and DPHON like '%f';
# Ե§ʣٽ
select WORD, K_F_FREQ from mrc2 where IRREG = 'Z' and TQ2 != 'Q' order by K_F_FREQ desc;
# ߿Ūʰ̣ĸ
select distinct WORD, FAM from mrc2 where FAM > 600 and CONC > 600;
# 䤹
select distinct WORD, IMAG from mrc2 order by IMAG desc limit 30;
# ̣ͭפʸ
select distinct WORD, MEANC, MEANP from mrc2 order by MEANC + MEANP desc limit 30;
# ̾ưʤʻˤäƶѥΰۤʤ
select WORD, WTYPE, DPHON from mrc2 where VAR = 'O';
ؤʬǤϤ褯Τ줿Ѹθåǡ١Τ褦#1131. 2̾ưŵŪʶѥ ([2012-06-01-1]) ȡ#1132. ñʻ̤γ ([2012-06-02-1]) ǻȤ Amano ʸˤơ¸ߤΤäMRC Psycholinguistic Database ϡ150837줫ʤʸåǡ١Ǥ롥Ƹ˸ŪӿŪ26°ꤵƤꡤʣʾŬ礹ΥꥹȤñ˺ФȤǤΤħŪäؤμ¸ѤåꥹȤʤɤӤä˻Ȥ뤬ѥȤ߹碌Ǥϡưפ˸׳ؤθѤǤ
ѥϼ¤¿ˤ錄롥ʸǿλ˻ϤޤꡤΥѥ˴Ť٤ϰϤˤʤߤǽŪʻɸȤơ familiarity, concreteness, imageability, meaningfulness ʤɤꤵƤ롥ʻʤɤ쥫ƥϤƬά졤ϥեʤɤη֥ƥλǤ롥ȯ䶯ѥλˤбƤ롥Ȥ߹碌ˤäơ褽ΤȤǤΤǤϤʤȻפ碌̤Ǥ롥
ǡ١ȸץɤǤ뤬ץѥ뤹ʤݤ¿Τǡ־ΥեѤΤǤ롥2ĤΥեѰդƤꡤ줾쵡ǽϸꤵƤ뤬̾ӤˤϽʬ
Online search (answers limited to 5000 entries): ѥκ٤꤬ǽϷ̤5000ޤǤ˸¤롥
Online search (limited search capabilities): Ϸ̤ο¤ϤʤŪʥѥκ٤ֻȯΥѥľܻʤɡˤϤǤʤ
Amano, Shuichi. "Rhythmic Alternation and the Noun-Verb Stress Difference in English Disyllabic Words." ̾Ų¤̾Ų¤ݽûס 15 (2009): 83--90.
Powered by WinChalow1.0rc4 based on chalow