ε ([2018-01-03-1]) Ʊ٤ȥڥĹ˴ؤǡ⤦ʬϤƤߤʲϡ٥Υȥå100졤200졤500졤1000졤2000졤5000졤10000졤20000졤50000ˤĤơ줾͡1ʬ̿͡ʿ͡3ʬ̿ͤɽǤ롥ѸˡǤδåǡȤƤɤ
| Min. | 1st Qu. | Median | Mean | 3rd Qu. | Max. | |
|---|---|---|---|---|---|---|
| Top_100 | 1.0 | 2.0 | 3.0 | 3.1 | 4.0 | 5.0 |
| Top_200 | 1.00 | 3.00 | 4.00 | 3.77 | 4.00 | 10.00 |
| Top_500 | 1.000 | 4.000 | 4.000 | 4.498 | 5.000 | 10.000 |
| Top_1K | 1.000 | 4.000 | 5.000 | 4.968 | 6.000 | 15.000 |
| Top_2K | 1.000 | 4.000 | 5.000 | 5.406 | 7.000 | 15.000 |
| Top_5K | 1.000 | 5.000 | 6.000 | 6.014 | 7.000 | 16.000 |
| Top_10K | 1.000 | 5.000 | 6.000 | 6.488 | 8.000 | 16.000 |
| Top_20K | 1.000 | 5.000 | 7.000 | 6.954 | 8.000 | 17.000 |
| Top_50K | 1.000 | 6.000 | 7.000 | 7.622 | 9.000 | 20.000 |

ɸäܿŦǤϤʤѸɤ߽ԤˤľƤ뤳ȤȻפ롥#1091. ;١ѡ ([2012-04-22-1]) #1102. Zipf's law ȸοա ([2012-05-03-1]) ǤŦ褦ˡ褯ɤ߽ñΥڥûۤΨ褤ȹͤ뤫դˡ¿ɤ߽ʤñǤоĹƤǤ롥ñΥڥ˸¤餺ñβˤĤƤƱͤθѤƤȻפ롥
ޤѸˡˤƸ3ʸʾ֤ʤФʤʤȤ#2235. 3ʸ§ ([2015-06-10-1]) 롥ϵǽȤ٤Τƹ⤤ˤĤƤŬѤʤäơε§Ͼ嵭θΨȤؤŪ¦̤ĤȤ롥
ٸǤФۤɡΥڥʿŪûȤˡ1Ĥˡ٥Υȥå100졤1000졤10000ʤɤΥꥹȤ˴Ťʸ̤ñ夲Ȥ롥#2096. SUBTLEX-US Word Frequency List ([2015-01-22-1]) Ф٥Ѥơȥå100졤200졤500졤1000졤2000졤5000졤10000졤20000졤50000ˤĤĴȥå100ΥꥹȤˤĤƤεǥꥹȤǺܤƤ̤Ǥꡤʤˤ s, ll ʤɥѥλͤͳ褹Ȥܤָפ⤢뤬̤ˤϱƶڤܤʤ
ʲ˥դ̤ꡤ̤ǤʿͥǡϥHTMLȡˡȥå100Ķٸ췲Ǥ62.00%ޤǤ3ʸʲΥڥǤ롥3ʸʲγʲ3ʬΥӤޤǡˤȤȤ٤Ƥȡȥå200줫50000Ĵ̤ޤǡ41.50%, 24.60%, 17.00%, 12.65%, 8.06%, 6.01%, 4.55%, 3.20%ܸꤷƤ

ε#2875. Ѹäʬۤγʺ˷ȥĶǤߤ ([2017-03-11-1]) ˰³Ѹ٤γʺˤĤƹͤƤߤä˷ľŪ˳ʺǧǤɸȤơʺ1%ΥȤΤ롥кѳؤǤСȥޥԥƥⰦѤƤ֥ȥå͵ؤνפǤ롥ɤΤ餤ȤɸФ褤ѸäˤĤƸС٤ǥȥå1%뤽ۤ¿ʤˤäơΤΤɤΤ餤ΥƤ뤫ɸȤʤ롥
Ʊ褦ˡٿ81.5ۤɤŪϤ GSL θɽȡ1850ۤɤε祳ѥ#1424. CELEX2 ([2013-03-21-1]) ˴ŤɽǷƤߤȥå1%ȥȥå0.1%Ǥͤϡʲ̤ꡥ
| GSL | CELEX2 | |
|---|---|---|
| 1% | 47.05% | 69.36% |
| 0.1% | 14.60% | 43.57% |
#1103. GSL ˤ Zipf's law θڡ ([2012-05-04-1]) ǡGeneral Service List (GSL) κ2000;θɽѤơzipfs_law ΩͻҤ±餷ٽ̤ι⤤θ줬ιٸǤϤʤĶٸǤ뤳ȡǤʳ¿θ줬ʤ٤ٸǤȤȤǧ줿ΤȤϡѸʤơ餯ˤθäʬۤʿԶѹդǤꡤ礭ʤФĤȳʺħŤƤ뤳ȤΤǤ롥
Τ褦ʬۤγʺɽŪʻɸˡꥢηкѳؼԥˤʬۤʿ¬ɸȤ1936ǯ˹ͰƤ˷ (Gini's coefficient) 롥ͤϼ̤X˱äƺ鱦غǤ٤㤤줫⤤ؤȽ¤١٤ΥY˼äƤĤʤȡ餫ηα夬ζȤʤ롥Ķ (Lorenz curve) Ȥ٤Ƥθ줬Ʊ٤ǸȤˤϥĶ45٤α夬ľȤʤִʿפդˡüȤơ1ĤθΤߤ٤Τ٤Ƥͭ¾Τ٤Ƥθ줬٥ξˡִʿפȤʤꡤĶϺLȤʤ롥̤ϡĶϡ45٤α夬βˡθ̤Ȥ롥˷ϡѤȡ45٤α夬ľѤդȤľջѷΨȤɽ롥äơ0ʿ1ʿȤȤˤʤ롥
ơGSL ǡեǷ̡˷0.812ȽФĶȡʲΤ褦ˤʤ롥

餫ʿʬۤȤ롥ʤߤˡGSL ʥѥθɽȤȡ˥˷Ͼ夬㤨С1790줫ʤ륳ѥ#1424. CELEX2 ([2013-03-21-1]) ˴ŤǤϡ0.950 ȤޤͤФˡ
ͤޤǤˡ (122) ˵ä2010ǯννʺ˷Ĥȡܤ 0.336ꥫ 0.380꤬ 0.510ɤ 0.246 Ǥ롥äμҲˤʿʼҲǤ뤳Ȥʬ
Ρؿܷкѡ١ҡӡ2016ǯ
n-gram ϡפ䥳ѥؤˤŪʳǰʤǤʡ#2324. n-gram ([2015-09-07-1]), #956. COCA N-Gram Search ([2011-12-09-1]) ȡˡƥȤꤷƤ n-gram ġϥͥåȤ¾ˤߤƤ뤬ƴʰץġCGIǼƤߤХåɤ Perl ⥸塼 Text::Ngrams ѤƤ롥
ִܸá (basic vocabulary) ȤѸϡĴˤ͡ʵ˽Ф魯ε#2660. glottochronology ȴܸá ([2016-08-08-1]) Ǥ줿褦ˡ̸ˤƤ⡤̤ˤƤ⡤ܸäȤϲʤΤɤޤǤϰϤޤΤҴŪ뤳Ȥ
glottochronology ˷ȤؼԤϡȼ̸Ū̻Ūʴ顤ܸåꥹȤΤԽƤ㤨СʬʬǤ Swadesh (456--57) ϡʥꥹȤϺʤȤȤǧĤġ200줫ʤƤ롥ΰHymes ("Lexicostatistics" 6) ͳǷǤ褦
all, and, animal, ashes, at, back, bad, bark, because, belly, big, bird, bite, black, blood, blow, bone, breathe, burn, child, cloud, cold, come take, count, cut, day, die, dig, dirty, dog, drink, dry, dull, dust, ear, earth, eat, egg, eye, fall, far, fat-grease, father, fear, feather, few, fight, fire, fish, five, float, flow, flower, fly, fog, foot, four, freeze, fruit, give, good, grass, green, guts, hair, hand, he, head, hear, heart, heavy, here, hit, hold-take, how, hunt, husband, I, ice, if, in, kill, know, lake, laugh, leaf, leftside, leg, lie, live, liver, long, louse, man-male, many, meat-flesh, mother, mountain, mouth, name, narrow, near, neck, new, night, nose, not, old, one, other, person, play, pull, push, rain, red, right-correct, rightside, river, road, root, rope, rotten, rub, salt, sand, say, scratch, sea, see, seed, sew, sharp, short, sing, sit, skin, sky, sleep, small, smell, smoke, smooth, snake, snow, some, spit, split, squeeze, stab-pierce, stand, star, stick, stone, straight, suck, sun, swell, swim, tail, that, there, they, thick, thin, think, this, thou, three, throw, tie, tongue, tooth, tree, turn, two, vomit, walk, warm, wash, water, we, wet, what, when, where, white, who, wide, wife, wind, wing, wipe, with, woman, woods, worm, ye, year, yellow
ΰȼȤ߹碌ΤǤꡤθβФ뤳ȤˤʤäȤ괰ʴܸåꥹȤϺʤΤ顤餫θĴԤʤˡΰˤȤΤϡ1ĤˡǤϤ롥ʤSwadesh ǯ¬οǤΤѤΤϡӸ줿100ΥꥹȤǤꡤϡ#1128. glottochronology ([2012-05-29-1]) ǷǺܤ̤Ǥ100ꥹȤΤۤǯŪͭ⤤Ȥո⤢ (Hymes, "More" 341)ˡ
ܸäˤĤƤϡ#308. Ѹκѱñꥹȡ ([2010-03-01-1])#1101. Zipf's law ([2012-05-02-1])#1961. ܥ٥ơ ([2014-09-09-1])#1965. Ūʸǡ ([2014-09-13-1])#2625. ťΥɸ줫μѸ ([2016-07-04-1]) ʤɤεȡ
Swadesh, Morris. "Lexico-Statistic Dating of Prehistoric Ethnic Contacts: With Special Reference to North American Indians and Eskimos." Proceedings of the American Philosophical Society 96 (1952): 452--63.
Hymes, D. H. "Lexicostatistics So Far." Current Anthropology 1 (1960): 3--44.
Hymes, D. H. "More on Lexicostatistics." Current Anthropology 1 (1960): 338--45.
glottochronology ˡɾ Hymes (11) ϡ⤦1ͤοؼ Gleason ˤ1950ǯȤʤ顤ѲΨ˴ؤ3ĤˤĤƲ⤷Ƥ롥
(1) Every lexical item at every given time has a certain probability of change.
(2) This probability of change is variable, and is influenced by both linguistic and non-linguistic factors.
(3) There exist certain sets of largely independent vocabulary items in which the probability of change within the group is large relative to the variability of that probability of change.
glottochronology Ǥϡ(1) (2) 餫ȤƤHymes ä˽פȻŦΤ (3) βǤ롥ˤСäˤϤĤ췲Ĥꡤ췲Ūꤷΰٹ礤ɤṲ̄θ췲ŪǤꡤٹ礤ɤŪ礭ȤŪˤϤ "basic vocabulary" "non-basic vocabulary" ʤɤζ̤ǰƬˤƤ뤳Ȥϴְ㤤ʤɬ餫Ǥʤ "(non-)basic" ȤѸȤ鷺ˡŪ׳ŪʼˡǡäʬФǽƤ롥ºݤθڤˤϡ¿θθäˤĤĴ줾ˤĤĹ֤ˤ錄̻ŪʸѲΨᡤӤȤƻʺȤɬפǤꡤ˷ФȤΤǤϤʤڲǽϳݤƤȤפǤ롥
̡뤤ϸ̸ˤơܸ (basic vocabulary) ȤϲȤϡҴŪΤƳüԤˤȤäƤľŪʬΤǤϤ뤬ϰϤҴŪΤε#2659. glottochronology lexicostatistics ([2016-08-07-1]) Ǥ줿褦ˡܸäƱ˴Ϳ°Ȥ (1) Ū commonness (or frequency), (2) ̸Ū universality (of semantic reference), (3) ̻Ū (historical) persistence 3郎ƤƤꡤ餬ߤˤ褽شطˤ뤳ȤΤƤ롥3Ĥ°γơˤɤ٤νŤߤĤǽŪ˴ܸäꤹ٤ˤĤơä˹դϤʤ
glottochronology ˤȤäƤϡܸäȤϤޤǸǯ¬ꤹ뤿κǤϤ뤬षκõβǡܸäȤϲȤοˡξ¦̤뤳ȤˤʤäΤǤϤʤȤפ롥glottochronology Ȥʬ̤ˤĤƤ¿ȽʤƤβǷ깭ƤϤФܼŪǤꡤʿ˸ػŪʹ礭Ȥ
glottochronology ȴܸäˤĤƤϡHymes (32--33) ܤƤΤǡȡ
Hymes, D. H. "Lexicostatistics So Far." Current Anthropology 1 (1960): 3--44.
ؤʬȤƤθǯ (glottochronology) ȸ׳ (lexicostatistics) ϡФƱѤƤglottochronology ϻϼԤǤ Swadesh ϡξѸȤʬƤ롥伫Ȥ#1128. glottochronology ([2012-05-29-1]) εǡξԤϰۤʤȤΩglottochronology ʸǯءˤϡꥫθؼ Morris Swadesh (1909--67) Robert Lees (1922--65) ˤä1940ǯ˳줿̻ؤ1ʬǤ롥μˡ lexicostatistics ʸ׳ءˤȸƤФ롥פȽҤ٤ϡѸˤĤƹͤƤߤ
ؼԡҲؼԤ Hymes (4) Swadesh ˰͵ʤ顤ξѸζ̤Τ褦Ƥ롥
The terms "glottochronology" and "lexicostatistics" have often been used interchangeably. Recently several writers have proposed some sort of distinction between them . . . . I shall now distinguish them according to a suggestion by Swadesh.
Glottochronology is the study of rate of change in language, and the use of the rate for historical inference, especially for the estimation of time depths and the use of such time depths to provide a pattern of internal relationships within a language family. Lexicostatistics is the study of vocabulary statistically for historical inference. The contribution that has given rise to both terms is a glottochronologic method which is also lexicostatistic. Glottochronology based on rate of change in sectors of language other than vocabulary is conceivable, and lexicostatistic methods that do not involve rates of change or time exist . . . .
Lexicostatistics and glottochronology are thus best conceived as intersecting fields.
Ĥޤꡤglottochronology lexicostatistics ʪξԤνŤʤʬʤפˤǯ¬ꤹ礬ʬˤȤäƤǤ褯Τ줿ʬǤ뤫顤ξԤ¾ƱȤʤäƤȤȤSwadesh 80ǯФäߤǤϡlexicostatistics ϡŻҥѥȯŸˤǯ¬Ȥ̵طνⰷʬȤʤäƤꡤμϰϤϹäƤȤ
ǰѤ Hymes ʸϡˤ "basic vocabulary" ȤϲȤŪʪĤˤĤƿƤäƤꡤɤβͤ롥"basic vocabulary" ϡcommonness (or frequency), universality (of semantic reference), (historical) persistence Τ줫°뤤ϤȤ߹碌˴ŤΤȳͼƤ뤬ƱʸϤդεˤĤƤܤܸäˤĤƤϡ#1128. glottochronology ([2012-05-29-1]) #1965. Ūʸǡ ([2014-09-13-1]) εľܤ˰äۤ#308. Ѹκѱñꥹȡ ([2010-03-01-1])#1089. ȸ; ([2012-04-20-1])#1091. ;١ѡ ([2012-04-22-1])#1101. Zipf's law ([2012-05-02-1])#1497. taboo ŪȤʤͳ (2) ([2013-06-02-1])#1874. ٸθݼ ([2014-06-14-1])#1961. ܥ٥ơ ([2014-09-09-1])#1970. ¿٤شط ([2014-09-18-1]) ʤεǴͿ˿ƤΤǡ⻲Ȥ줿
Hymes, D. H. "Lexicostatistics So Far." Current Anthropology 1 (1960): 3--44.
#53. Ѹ through ֤515̤ ([2009-06-20-1])#219. eyes ɽ172ֻ̤ ([2009-12-02-1]) ˰³ѸˤˤֻѰۤˤĤơϡֻμ¿ɾΤʡ "such" 夲롥
#1622. eLALME ([2013-10-05-1]) ǾҲ𤷤ѸϿ LALME βŻ eLALME ˤơItem List 10 "such" äƤ롥ΰֻȴФȡԳΤƾܤ˿Ƥ⡤ʲ134बʤäοͤʸڤ١ˡ
asoche (1), aswyche (1), schch (1), schech (1), scheche (3), schiche (1), schoche (1), scht (1), schuc (1), schuch (3), schuche (4), schut (1), schute (1), sclik (2), sclike (1), sclyk (2), sclyke (2), scoche (1), scwche (1), sech (8), seche (39), sewyche (2), shich (1), shiche (1), shoch (1), shoche (1), shuch (5), shuche (3), shych (1), sic (6), sic- (1), sich (53), siche (101), sick (1), sɩͨh (1), sik (1), sik- (1), sike (2), silk (3), sli (1), slieke (1), slik (10), slike (26), slilk (2), slkyke (1), slyk (13), slyke (26), soch (12), soche (60), souche (3), sowche (2), soyche (1), squike (1), squilk (2), squylk (1), sqwych (1), sqwyche (1), sswiche (1), suc (1), succh (1), sucche (5), such (242), suche (375), suchee (1), sucheȝ (1), suchet (1), sucht (1), suchte (1), suech (4), sueche (6), suhc (1), suhe (1), suich (9), suiche (7), suilk (6), suilk- (1), suilke (3), suilkin (1), sulc (1), sulk (4), sulke (2), sutche (1), suth (1), suuch (1), suuche (1), suuech (1), suueche (1), suych (13), suyche (15), suylk (7), suylke (6), svche (1), sviche (1), swc (1), swch (7), swche (4), swech (19), sweche (48), swelk (4), swhiche (2), swhilke (2), swhych (2), swhyche (1), swic (2), swich (77), swiche (84), swichee (1), swilc (3), swilk (76), swilke (45), swilkes (1), swisɩͨhe (1), swlk (1), swlke (1), swuch (3), swuche (2), swych (56), swyche (65), swyeche (1), swyk (1), swyke (1), swyl (1), swylk (62), swylke (35), swylle (1), syc- (1), sych (23), syche (67), syge (1), syk (4), syk- (1), syke (5), sylk (3), sylke (2)
̤ٳ뤷٤פȡȥå10 suche, such, siche, swiche, swich, swilk, syche, swyche, swylk, soche Ǥ롥ȥåפ2 suche such ϸѸƤԤˤȤäơʬQŪˤߤºݡ2617㤬ʸڤ졤1867Τۤ3ʬ1롥ޤȥåפ10ǡۤ3ʬ2롥äơֻ¿뤫ȤäơΤޤʤ뺮٤ȤȤˤϤʤʤΤ褦ʻϡѸ¿ֻǧ¿θˤĤǧ졤٤Τʤˤ⤢٤餷ΤɤäƤȤ롥ȤƤ⡤νɤˤȤäƤϤϤؤʾä˰㤤ʤˤĤƤϡ#1311. ֻɸಽϤʤɬפ ([2012-11-28-1])#1450. Ѹֻ¿ϤϤؤǤ ([2013-04-16-1]) ̤Ǥ롥
ѸḽνĴ٤СäȰֻμŵϼǰˤΤΡ500ۤɤȤȤ롦
ε#2362. haplology ([2015-10-15-1]) ǥꥷ haplo- (one, single) ˿줿θ캬˴ϢƤ⤦1ʸؤ伭ؤѸȤƤФнв hapax (legomenon) 夲褦ΤʤǡʥǤϤʤȡǡ1٤ѤƤʤʶˤؤꥷ hapax (once) + legomenon (something said) ʤʣʣ hapax legomena Ȥ
"nonce word" hapax legomenon ƱȤƤ뼭⤢뤬Ԥϡפ֤λ¤Ѥפؤnonce-word ϿŪǰƬѤ뤳Ȥ¿ΤФhapax legomenon ʸ˸1٤Ǥ뤳Ȥ˾ƤƤȤ㤤롥nonce ʤξ¤ΡˤȤθ츻ˤĤƤϡ#1306. for the nonce ([2012-11-23-1]) ȡ
hapax legomenon ϡȤδϢǡФиڤƤˤ롥OED ˤȱѸˤ1692ǯΤȤǡ"J. Dunton Young-students-libr. 242/1 There are many words but once used in Scripture, especially in such a sence, and are called the Apax legomena." Ȥ롥
ʸؤ츻ؤˤơhapax legomenon ϤФȤʤ롥θθ츻Ϥ̣Ǥ뤳Ȥʤʤ伭ؤǤϡΡָפȤǧƤ褤Τδְ㤤ǤϤʤ˷Ǻܤ٤ݤȤƬˤ꤬ (see #912. ʤ (3) ([2011-10-26-1])) ǡ䤽Ȥϡhapax legomenon ϽפʹͻоݤȤʤ롥ȤΤϡ1٤Ū˽и뤿ˤϡüԤŪʸȤʤФʤʤǤ (see #938. (4) ([2011-11-21-1]))
ºݤΤȤ halax legomenon Ϸ褷ƾʤʤΤȤϡåפˡ§˾Ȥ餻жä٤ȤǤϤʤ (see #1101. Zipf's law ([2012-05-02-1]), #1103. GSL ˤ Zipf's law θڡ ([2012-05-04-1])) ѸȤƤϡChaucer Ѥnortelrye (education) Shakespeare honorificabilitudinitatibus, ޤ Dickens sassigassity (audacity?) ʤɤ롥
伫ʬѤ n-gram Ȥʬϼˡ롥ѥؤǤ⤹ǤˤߤγǰǤꡤɽ (collocation) θʤɤǤΤ褦Ѥ褦ˤʤäΥѥΥեˤƤѤƤꡤ#607. Google Books Ngram Viewer ([2010-12-25-1]) Ǥ̾˴ޤޤƤۤɤܥ֥Ǥ COCA (Corpus of Contemporary American English) N-gram ǡ١Ѥơ#956. COCA N-Gram Search ([2011-12-09-1]) ƤʤαѤϡ#953. ƬƧ2।ǥ ([2011-12-06-1])#954. ӱƧ2।ǥ ([2011-12-07-1])#955. ʸϤ碌2।ǥ ([2011-12-08-1]) ȡˡBNC ǤϡExplore Words and Phrases from the BNC ѤǤ롥
ԥ塼ѤʬϼˡȤʹ뤬n-gram ιͤϻäñǤ롥ʸ٥ 2-gram (bigram) ͤƤߤ褦ĹαñȤ pneumonoultramicroscopicsilicovolcanoconiosis (#63. پɤϱѸǺǤŤµ ([2009-06-30-1])) ˤȤ롥ޤƬ2ʸ1Ȥ pn Фˡ2ʸܤ˿ʤƱ褦 ne Ф3ʸܤ˿ʤ eu 4ʸܤ˿ʤ um 롥Ʊ褦ˡ1ʸıˤ餷ʤ顤Ǹ is ޤ2ʸ1ȤȽäƤ44Ȥ2ʸȤˤʤ롥ȤΤʤǡic co ȤȤ߹碌ϳơ3ꡤos, si, no, on Ȥ߹碌ϳơ2졤ʳȤ߹碌Ϥ1٤Ǥ롥äơñˤƺǹ٤2ʸ1Ȥ ic co Ȥʤ롥
n-gram ñ̤ϡΤ褦ʸǤɬפϤʤǤǤ褤礭ñ̤ǤǤǤ褯礭ʶʤɤΤ礭ñ̤Ǥ褤Ѹ쥳ѥؤǤϡȤñ̤ǹͤΤ̤Martin Luther King, Jr. I Have a Dream αΥƥȤǸñ̤ 4-gram ȡǤ¿4Ȥ߹碌ϡͽ̤ "I have a dream" 8"will be able to" Ʊ8롥"Let freedom ring from" 7Ȥ褯롤ʬϤǽȤʤ롥Ǥ4ȤפꤷΤ 4-gram ȸƤФ뤬ܤ뤤Ĥʸθ뤫ˤ 1-gram (unigram), 2-gram (bigram), 3-gram (trigram) 5-gram ʾͤ뤳ȤǤ1-gram ξ硤ꥹȤϡ¾ƸɽǤˡ
祳ѥ줿 2-gram 3-gram ΰϡ켫ΤɽθʤɤǤϴܥǡȤʤ뤿ᡤ־Ǥ⤤ȸƤ롥ܸǤN-gram ѥ - ܸ쥦֥ѥ 2010뤷ѸǤ COCA n-gram ǡ١ 롥ޤBigram Plus Ǥϡ˱Ѹ쥳ѥޤƼѸ쥳ѥ N-Gram Search Ԥʤ뵡ǽƤ롥ۤˤǤդΥƥȤ䥳ѥоݤ n-gram ƼΥġ䥽եȤ⡤־ǽ
n-gram ʬϤθʬؤαϰϤϹʲʸˤϲȤͽ¬ǽȤطᡤˤ벻ǧʬϡȽꡤưڥåõΡʸѥǥåκʤɤ˳Ѥ롥ɽθǤϡܤˤԲķμʤȤʤäƤ롥n-gram ϤäѤȤɽ̲줿ƥȤоݤȤؤˤ빽¤ˤޤä뤳ȤʤᡤʸˡΤ褦ʸ̤Ƚ褦ܤϡn-gram in Wikipedia ȡ
n-gram ϹǡޤޤȤƻꤽ˱ѸƥȤˤ⡤ѤƤ
ʸ嵭2015/09/12(Sat) Sketch Engine N-grams ⻲ȡ
Ѳ٤δطˤĤƤϡ#1239. Frequency Actuation Hypothesis ([2012-09-17-1])#1243. ٤θ̻ŪΤˡ ([2012-09-21-1])#1265. ٤ȲѲνδط˵ŤƤ Schuchardt ([2012-10-13-1])#1864. ȴդٸ̡ ([2014-06-04-1]) ʤɤǵƤεǤϡٸٸǤϤɤ餬˸Ѳ˴ޤ뤫ʤɤνꡤ뤤ϥ塼 (schedule_of_language_change) ˼㤬äȤ̤ˡٸ줢뤤ٸˤʤѲǤȤष䤹ѲǤȤΤ褦ʤȤϤΤ
Ѥ Fortson (659) ϡʸˡ (grammaticalisation) ̣Ѳ (semantic_change) ٤δطˤĤƤ뤬ȤƤξԤδ֤شطϤʤȽҤ٤Ƥ롥Ĥޤꡤʸˡ̣ѲϹٸˤѤ뤷ٸˤƱ褦˺Ѥ롥ˡ٤ȤװƳɬפϤʤȤ
We have seen, then, that both frequent and infrequent forms can be reanalyzed; both frequent and infrequent forms can be grammaticalized. If all these things happen, then frequency loses much or all of its force as an explanatory tool or condition of semantic change and grammaticalization. The reasons are not surprising, and underscore the sources of semantic change again. Frequent exposure to an irregular morpheme, for example (such as English is, are), can insure the acquisition of that morpheme because it is a discrete physical entity whose form is not in doubt to a child. By contrast, no matter how frequent a word is, its semantic representation always has to be inferred. Classical Chinese shì was a demonstrative pronoun that was subsequently reanalyzed as a copula; exposure to shì must have been very frequent to language learners, but so must have been the chances for reanalysis.
Fortson Τ褦˼ĥΤϡε#2175. Ūʰ̣ѲؤȽ ([2015-04-11-1]) Ǥ줿褦ˡ̣ѲˤϡФпƤ褦Ϣ³ϤʤषϢ³ŪʤΤǤȹͤƤ뤫Ǥϡʸˡ̣Ѳ¾θѲۤʤ櫓ǤϤʤٸ줢뤤ٸɤȤǤϤʤȤ٤ͿƤ뤫Τ褦˸ȤС diffusion (or transition) μˤƤǤäơimplementation μˤ٤δͿϤʤȡϡ#1872. Constant Rate Hypothesis ([2014-06-12-1]) ۵ѲѤǤ롥
Fortson IV, Benjamin W. "An Approach to Semantic Change." Chapter 21 of The Handbook of Historical Linguistics. Ed. Brian D. Joseph and Richard D. Janda. Blackwell, 2003. 648--66.
ɸˤĤơ#1620. Ѹˤ /t, d/ źá ([2013-10-03-1]) #1575. -st θźä˴ؤ Dobson ιͻ ([2013-08-19-1]) εǼ夲Ƥä˸ -st ˤ t οˤĤƤ ##508,509,510,739,1389,1393,1394,1399,1554,1555,1573,1574,1637,1807,2062 γƵˤƤ
Ѹˤ /t/ æޤȹϤ¸ߤ뤬Wełna (329--30) OED MED 齸ΰͿƤƤΤǡǺܤWełna ϡ/t/ æ "permanent" ʤΡʸѸޤǤθ̤³ƤΡˤ "sporadic" ʤΡʰθ̤줿ˤȤη֤زΡˤȤ̤줫Ѹ줫ǶʬƤ롥
(1)
(a) Permanent t-insertion in native words: ME behest (<OE behæs); against, amidst, amongst, betwixt
(b) Permanent t-insertion in foreign words: ME ancient (<ME auncien), ME cormorant (<F cormoran), ME ernest) (<ME ernesse) 'earnest' (=pledge money), ME pagent (<ME pagyn 'pageant', ME perchement (<ME parchemin) 'parchment', ME fesaunt (<F fesan) 'pheasant', ME truant (<F truan), ME tirant '<F tiran) 'tyrant'
(c) Sporadic t-insertion followed by t-loss: ME glisten (<OE glisnian, ME listen (ONhb. lysna); ME vermin (<ME vermint <F vermin
. . . .
(2)
(a) Permanent t-loss in native words: (a) anduel (< onfilt) 'anvil'; ME best(a) (<betsta), ME blesse (<bletsen), OE blosma (<blostma), ME last(e) (<lattste); ENE bussle (<bustle), ENE brisle (<ME bristle), ENE miscelto (<ME mistilto) 'mistletoe', ME nestle, ME ?rustle, LME thrissil (<OE þistil) 'thistle', ENE throssle (<ME þrostle), Sc. quhissle (<OE hwistle) 'whistle', ENE wressel (<ME wrestlen); christen (OE cristnian <Lat.), ME fasten (<OE fæstnian); ME offen (<ME often).
(b) Permanent t-loss in foreign words: (a) ME apostle, castle, epistle, ENE iussell (<LME iustil) 'jostle', LME pestle, ME tresselle (<trestle) 'trestle' (obs.); crysmas (<Cristmasse) 'Christmas'; (b) ME chasten, ENE chestnutte (<chest-nut), ENE hasten; (c) ENE craven (<ME cravant), ME orisoun (<ME orizonte) 'horizon'
(c) Sporadic t-loss: ENE paisan (<OF paysant) 'peasant'.
Wełna Ѹ줫Ѹ t οĴ(1) ǤϹٸ줬ѸǤٸ줬βѲαƶ䤹(2) ʬۤ餫Ϥߤʤ(3) t ϼȤѸθݤǤ t æϼȤƽѸθݤǤ뤳ȡ3ȤƼƤ롥ɬǤϤʤ⤤Ĥ롥ܺ٤Ĵ˾ޤ롥
Wełna, Jerzy. "Insertion and Loss of the Voiceless Dental Plosive [t] in Middle English." Studies in Middle English: Words, Forms, Senses and Texts. Ed. Michael Bilynsky. Frankfurt am Main: Peter Lang, 2014. 329--42.
ǧθؤθѲ˴ؤǥȤơѴץǥ (usage-based model) ȤΤƤƤ롥ëˤȿ (106, 105) 狼䤹
뤳ȤФˡζȤʤ륹 [A] 顤餫ǰæȤäˡ (B) 롥Ϥᡤ(B) ϥ [A] ˹פʤ(B) ˡ֤夹ˤĤơ(B) [A] ȶˤθΥƥ˼ޤ褦ˤʤ롥ȡ(B) Ǥ餿ʥ [A'] Ф졤ˤä (B) ǧ褦ˤʤäƤΤǤ롥Τ褦ѲΥƥֻѴץǥפ뤤ϡˡץǥ (usage-based model) Ȥ (Langacker 2000) ë106

ޤϽФϡݲǤȤǡʸˡ§ϽФȤӤ롥̾ʸˡ§ŪǤΤФơޤưŪǤꡤǤȤ㤤롥ޤϡæ㤬夹ˤĤơѹƤޤѲβˤơæ㤬夹ٹ礤ˤϸĿͺ뤿ᡤɬŪ˥ΤθĿͺ뤳Ȥˤʤ롥ѲΤ褦˰֤ŤƤȤ館ѴץǥˤƤϡηϤΤΤήưŪʤΤˤߤ
ޤ٤˸ĿͺȤȤϡѲ® (speed_of_change) ľ뤹뤷θλ (frequency) 䶦 (collocation) ȤϢѴץǥϡδطˤܤƤ롥ѲʥߥåʤΤǤϤ뤬줽ΤΤ˥ʥߥåʤΤǤꡤΥʥߥθλѤΤʤˤȤȤƶĴɾǤ
ë سؤӤΥǧθء١ҤĤ˼2006ǯ
αѸظˤơҤɽȤХꥫѸ˴ؤ Kucera and Francis (1967) ΤΤ䡤ꥹѸŤ֤꿷ΤȤ CELEX (1993) 䤽2 (cf. #1424. CELEX2 ([2013-03-21-1])) 褯ѤƤǶᡤȽˡ˴ŤꥫѸθɽ줿٥륮إؤμ¸زʤ SUBTLEXus Ǥ롥HP顤SUBTLEXus ΰ췲ɽΥե䵭ҤɡǤ롥
SUBTLEXus δפˤ륳ѥϡ8388αDzλνǤꡤ5100˵ڤ֡SUBTLEXus ɽϡKucera and Francis CELEX ɽ٤ơĤλФ줿ɸˤƤƤȼĥƤ롥٤ϡФ (lemma) ȤǤϤʤ (word form) Ȥ˿Ƥꡤ㤨̾Ǥñ -s ʤɤʣ̰ʰۤʤ74,286ˡ̾ưʤʣʻȤѤˤĤƤϡ줾ʻ줴Ȥ٤ˤ⥢Ǥ뤷ͥʻ (Dominant POS) Τۤع绻٤ؤ⥢Ǥ롥ǡˤϡۤ˲αDz˸Ƥ뤫ʸȤƸƤΤϲ٤пäɸZipf ɸ (cf. #1101. Zipf's law ([2012-05-02-1])) ʤɤޤޤƤ롥μΥǡޤޤƤȡŪȥǥǤͭѤǤäե١Ǥ뤳Ȥ⸲ħ
ɤǤ뤤ĤΥǡΤʤ "a zipped Excel file of SUBTLEX-US with the Zipf values included" ɤäƤߤ㤨С(1) Ū¿졤 (2) ¿αDzˤ⸽ϡŪʰ̣٤⤤ȹͤ (1) (2) ˴ؤпλɸݤ碌ơ߽¤٤ƺǽ100ȡäκñ100줬ϤάҳʤɤޤޤƤ뤬ʲΥꥹȤǤ롥
you, I, the, to, s, a, it, t, that, and, of, what, in, me, is, we, this, he, on, for, my, m, your, don, have, do, re, no, be, know, was, not, can, are, all, with, just, get, here, but, there, ll, so, they, like, right, out, go, up, about, she, if, him, got, at, now, come, oh, one, how, well, want, yeah, her, think, good, see, let, did, why, who, as, going, his, will, from, when, back, time, yes, look, d, take, an, where, man, would, them, been, some, or, tell, us, had, were, say, could, gonna, didn, hey
ۤˤϡ10졤25졤50졤100졤250졤500졤1,000졤2,500졤5,000졤10,000졤25,000졤50,000졤100,000ˤĤơDominant POS Ȥ˿夲Ƥߤ뤳Ȥ⤿䤹#666. COCA 5000ʻ̤γϡ ([2011-02-22-1])#667. COCA 50ʻ̤γϡ ([2011-02-23-1])#1132. ñʻ̤γ ([2012-06-02-1]) εǤ⡤̤Υѥˤ褦ĴԤäSUBTLEX-US ǤĴ̤ϼΥդˤޤȤ롥
ʲϤޤθġ (SUBTLEX-US Word Frequency Extractor) ޤʤΤǡ10ޤǤ̤ϤʤͤǤSUBTLEXus ʣʸǽʡSUBTLEXus Online Search ɤ
ܸäȸƤФΤ¿ŪˤĤơ#1960. ѸäΥԥߥåɹ¤ ([2014-09-08-1]) ǿ줿ܸäȤϡŪ٤⤯˽졤Ѳˤ̣ˡ¿ˤ錄ʤɤħġϢϡ#308. Ѹκѱñꥹȡ ([2010-03-01-1])#1089. ȸ; ([2012-04-20-1])#1091. ;١ѡ ([2012-04-22-1])#1101. Zipf's law ([2012-05-02-1])#1874. ٸθݼ ([2014-06-14-1])#1961. ܥ٥ ([2014-09-09-1])#1965. Ūʸǡ ([2014-09-13-1]) ¾εǤȰäƤ
ϡȴϢơٸ¿ŪǤȤ̿ˤĤƹͤƤߤ٤ι⤤ۤɸ¿٤㤤ϸ¿⤿ʤȤȤϸѤλ¤˾Ȥ餷Ƽ¾ڤޤŪˤZipf's law Τ Zipf ϡΩ줫餳βĩ
Zipf ϡE. L. Thorndike αѸ20,000 Thorndike-Century Senior Dictionary ˴Ť٤ȸشطõäμϡŸѸʤɤü register ĸϷǺܤƤ餺ŪѤΤߤǺܤƤ롥ðǰĴٰ̡ȡ°줬ʿѸȤδ֤ˡ餫شط줿ʲϡZipf (253) ˼Ƥ륰դƸΤǤ롥ξȤпǤꡤXٽ̤YٰʿѸɽ魯

Ϥۤ0.5üԤȯäİԤİˤѤ˴ؤͽ¬礹ȤοŪդϻĶΤDzǤʤZipf ϷȤƸθ١ʽ̡ˤδطˤĤƼΤ褦꼰 (Zipf 255)
. . . different meanings of a word will tend to be equal to the square root of its relative frequencies (with the possible exception of the few dozen most frequent words)
طʤˤϡ¿䤢θ˶ʬ뤫Ȥạ̈¦䤦٤⤪ˤ뤬٤Ǥ롥Ϣơ#1091. ;١ѡ ([2012-04-22-1]) zipfs_law γƵ⻲Ȥ줿
Zipf, G. K. "The Meaning-Frequency Relationship of Words." Journal of General Psychology 33 (1945): 251--66.
#1872. Constant Rate Hypothesis ([2014-06-12-1]) Kroch ξѲΥ塼˴ؤ벾֤Ϥ褽Sɽ蘆졤ΥѥϰۤʤŪĶˤƤ֤ⷫȤۤʤĶˤƤ⡤ѲƱߥƱѲΨǿʹԤȤΤβǤ롥⤷ĶȤζѥǤϤʤ褦˸ˤϡϴĶȤ˿Ƥ٤ǽŪʸŪװˤۤʤ뤫Ǥ롤Ȳ᤹롥
Kroch ΤβΩ벾ФƤ롥1Ĥϡ#1811. "The later a change begins, the sharper its slope becomes." ([2014-04-12-1]) ǾҲ𤷤Ǥ롥Ѳɽ魯ϡĶȤ˥ѥǤϤʤĶȤ˳ϻۤʤ뤷ѲΨ®١ˤۤʤ롤Ȥͤβϡ˰ʤǡϤĶǤѲŪˤäʹԤ뤬٤ϤĶǤѲŪ˵®˿ʹԤ롤ʤٵޡפĥ롥伫Ȥ⤳ε˴Ť路ʸ (Hotta 2010, 2012)
̤β⤢롥ѲɽĶȤ˥ѥǤϤʤĶȤ˳ϻۤʤ뤷ѲΨ®١ˤۤʤ롤ȼĥǤϾҤΡٵޡפβƱषȵդΥ塼Τ롥ĤޤꡤϤĶǤѲŪ˵®˿ʹԤ뤬٤ϤĶǤѲŪˤäȿʹԤ롤ȡٵޡפʤ̡ٴˡפǤ롥ϡBailey ѲθȤƼĥƤΤ1ĤǤ롥
Bailey ϸѲΥ塼˴ؤơ2ĤθƤ롥1ĤϸѲSȤΡ⤦1ĤϾ嵭θѲΡٴˡפȤĥ줾졤ĥսѤ褦
A given change begins quite gradually; after reaching a certain point (say, twenty per cent), it picks up momentum and proceeds at a much faster rate; and finally tails off slowly before reaching completion. The result is an ʃ-curve: the statistical differences among isolects in the middle relative times of the change will be greater than the statistical differences among the early and late isolects. (77)
What is quantitatively less is slower and later; what is more is earlier and faster. (If environment a is heavier-weighted than b, and if b is heavier than c, then: a > b > c.) (82)
2ܤΡٴˡפˤĤƤϡΰѤʬȤꡤŪ¿ʤüŪˤ١ˤȤѥͿƤ뤷Ƥ뤳Ȥդ줿ɤΤȤѲΥ塼뤬ĶȤ˰ۤʤ뤫ȤϡѲΥ塼⤦1Ĥ礭ꡤʤ٤ȸѲνȤȤͿʤΤ⤷ʤȻפ碌롥
ǤϡɤβΤ٤ꤹ뤳ȤϤǤʤ̤θѲˤĤơ¤иŪ˽ƤʤΤ
Kroch, Anthony S. "Reflexes of Grammar in Patterns of Language Change." Language Variation and Change 1 (1989): 199--244.
Hotta, Ryuichi. "Leaders and Laggers of Language Change: Nominal Plural Forms in -s in Early Middle English." Journal of the Institute of Cultural Science (The 30th Anniversary Issue II) 68 (2010): 1--17.
Hotta, Ryuichi. "The Order and Schedule of Nominal Plural Formation Transfer in Three Southern Dialects of Early Middle English." English Historical Linguistics 2010: Selected Papers from the Sixteenth International Conference on English Historical Linguistics (ICEHL 16), Pécs, 22--27 August 2010. Ed. Irén Hegedüs and Alexandra Fodor. Amsterdam: John Benjamins, 2012. 94--113.
Bailey, Charles-James. Variation and Linguistic Theory. Washington, DC: Center for Applied Linguistics, 1973.
ٸϷŪݼŪȤȤϡ褯롥٤Ѳ䤹Ȥδ֤شط餷ȤϡǶǤ#1864. ȴդٸ̡ ([2014-06-04-1]) Ǽ夲εƬˤ⤤ĤεؤΥĥä (##694,1091,1239,1242,1243,1265,1286,1287,1864) ٸϤƤŪʴøǤ⤢뤳Ȥ顤ϴøäݼȤˤ̤롥ºݡ#1128. glottochronology ([2012-05-29-1]) ϡøäݼȤä
Ȥ (sign) ֤Ȥɽ (signifiant) ˤĤ٤ݼδطŦΤʤСΤ⤦1Ĥ¦̤Ǥ뵭 (signifié)ʤ̣ˤĤƤƱͤδطŦƤ褤Ϥ٤ΰ̣ǤСѲˤȤΤǤϤʤStern (185) ϡٸθݼ˿Ƥ롥
It is well known that the most common words of a language retain most tenaciously old and otherwise discarded forms and inflections. It is reasonable to assume that a strong tradition has similar effects on meanings. Note, however, that the retention of one or more old meanings is no obstacle to the acquisition of new ones: frequency is only a conservative factor for already established meanings.
٤θĸϤ켫Τ٤ǤꡤŪʴøǤΨ⤯ΤȤݼŪȤͽۤ롥ʤۤ Stern νҤ٤̤ꡤι٤θݼŪȤƤ⡤θŪŪʸղä뤳Ȥ˸櫓ǤϤʤष¿ξ硤ø¿Ǥ롥
Stern θ褦ˡ֤ˤĤƤ̣ˤĤƤ⡤٤ݼȤشطǧ褦˻פ롥1礭ʰ㤤롥֤Ѳˤϡ˼ä뤤ϾʤȤξ variants Ȥ¤ΩľȤ롥ΤȤ ơ variant Τ˶̤롥̣ѲˤϡŤɬ֤ΤǤϤʤξѤƤ椯Ȥ¿ variants ִΩȤϡ餬¿ȤѤ߽ŤʤäƤǤ롥֤ݼỌ̇̄ݼϡΰ㤤ռʤƤɬפϢơ#1692. ̣ȷ֤δط ([2013-12-14-1]) 3ѤȤ줿
Stern, Gustaf. Meaning and Change of Meaning. Bloomington: Indiana UP, 1931.
Ѳˤܤ (frequency) ˤĤơfrequency γƵȤ櫓#694. ٸԵ§ʣ ([2011-03-22-1])#1091. ;١ѡ ([2012-04-22-1])#1239. Frequency Actuation Hypothesis ([2012-09-17-1])#1242. -ate ưζܹԡ ([2012-09-20-1])#1243. ٤θ̻ŪΤˡ ([2012-09-21-1])#1265. ٤ȲѲνδط˵ŤƤ Schuchardt ([2012-10-13-1])#1286. ֲѲΰۤʤ2ưŤ ([2012-11-03-1])#1287. ưζܹԤ١ ([2012-11-04-1]) ͡˵Ƥ
ŪʰȤƤ Phillips ˤ#1239. Frequency Actuation Hypothesis ([2012-09-17-1]) ˴ؿƤ뤬̸ܸˤΤȴ (innovative potential) γȻ٤δĴʸĤƤ "Revised Frequency Hypothesis of Analogical Leveling (RFH)" ܤ줿
ʸǡԤ Matsuda ϤޤǤ˻ŦƤ͡ʸŪҲŪʸŪװ˴ͿŪŪ餫ˤ褦ȤθƤѿϰʲ̤Ǥ (12)
I. Linguistic
1. Length of the stem [measured in mora]
2. Conjugation type of the verb: i-stem/e-stem
3. Conjugation form following the potential suffix: Negative/Others
4. Morphological structure of the preceding stem: Monomorphemic verb/Compound verb/Auxiliary verb/Causative verb
5. Type of clause in which the potential form is embedded:
Main clause
Semi-embedded clause (Adverbial clause/Gerund)
Embedded clause (Quote/Relative clause/Predicate complement clause/Noun complement clause)
II. Social
1. Age
2. Sex
3. Area of residence: Uptown/Downtown
III. Style (taken from Labov & Sankoff, 1988)
Casual (narrative, group, kids, tangent)
Careful (response, language, soapbox, careful)
Matsuda ϤѿʤȤ߹碌ˤθܤŪ˳ΤƤͳơȴդγ˴ͿŪǤ뤳ȤSex Area of residence ʤͭպνФʤѿ⤢äơʬϤΤȤǡθǤϹθƤʤä٤ȤѿƳMatsuda ϡٸ̤ȤΤͤ褦ȤȤˡä٤θФ褤ΤȤܼŪ˸ڤƤ롥㤨 mirare(ru)/mire(ru) ʸʤˡʤˡˤξˤϡ촴 mir- ٤٤Τ뤤ܼ -are/-e ٤٤ʤΤԤǤСmir- Υȡ٤ȥ٤Τɤˤ٤ʤΤѸ drive--drove ʤԵ§Ѳư˸ݤ٤δʬϤˤϡޤΤΤ٤ˤФ褵Ūܸ mir-e(-ru) ʸʤˡˤξˤϡɤηǤ٤Ф褤Τ
ʾΤ褦ʹͻФơMatsuda ϶巿ȤǤϹθ٤٤ñ̤ۤʤäƤΤǤϤʤȤ롥줬Ҥ "Revised Frequency Hypothesis of Analogical Leveling (RFH)" (24) Ǥ롥
In analogical leveling, the token frequency of the unit undergoing the leveling and its degree/rate of leveling tend to show an inverse correlation, where the "unit" is defined according to the degree of fusion of the form undergoing the leveling with its neighboring morpheme(s). If the form is highly fused with the neighboring morpheme, the whole (morpheme, form) combination counts as a "unit" whose frequency is to be measured. If it is not, the form alone counts as a "unit," and its own frequency suffices as a correlate of the rate of leveling.
βѤСΤȴդĴη̵̤ʤǤȤĤޤꡤȴդγȻФơȿŪˡ˴ͿŪ٤Ȥϡβǽɽ魯ܼ -are/-e 켫ΤΥȡ٤Ǥꡤܤư촴Υȡ٤䥿٤ǤϤʤȡ
⤷Τ褦ˤפ뤬ѤʵȤơ٤㤤ܼȤäơʤ켫ȤοʿʤȴˤʤळȤˤʤΤ-are ˤƤ -e ˤƤ٤㤤ΤǤСʤԤԤִƤ椯ȤˤʤΤ
Matsuda, Kenjiro. "Dissecting Analogical Leveling Quantitatively: The Case of the Innovative Potential Suffix in Tokyo Japanese." Language Variation and Change 5 (1993): 1--34.
ǯΤȤˤʤ뤬Ƹ쥳ѥ ARCHER: A Representative Corpus of Historical English Registers Untagged 줿ܺ٤ϡ Documentation뤤 VARIENG ˤѥβɤѸ˸Υ饤ꡤߤεARCHERοǸ⻲ͤˤʤ롥
ARCHER ϡ1990ǯƬ Biber and Finegan ԻƤΤǡߤǤ14ؤƱǴƤ롥2013ǯ˸줿3.2Ǥ Manchester ( David Denison and Nuria Yáñez-Bouza) ˤǤ롥ѥƤӤüŪɽС"a multi-genre historical corpus of British and American English covering the period 1600--1999. The corpus has been designed as a tool for the analysis of language change and variation in a range of written and speech-based registers of English." ȤȤǤ롥
ѥεϤ1,710ե롤3,298,080줫ʤꡤǤα6:4ۤɡޤȤ8Ƥˤ12˥ʬƤ (a = advertising, d = drama, f = fiction, h = sermons, j = journals, l = legal, m = medicine, n = news, p = early prose, s = science, x = letters, y = diaries) եȸϰʲ̤ꡥ
| BRITISH | a | d | f | h | j | l | m | n | p | s | x | y | TOTAL | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1600--49 | files | 0 | 10 | 0 | 0 | 0 | 10 | 0 | 0 | 10 | 0 | 0 | 0 | 30 |
| words | 0 | 32,342 | 0 | 0 | 0 | 21,026 | 0 | 0 | 32,741 | 0 | 0 | 0 | 86,109 | |
| 1650--99 | files | 0 | 10 | 11 | 10 | 10 | 10 | 21 | 10 | 0 | 10 | 75 | 10 | 177 |
| words | 0 | 30,328 | 41,667 | 21,818 | 21,186 | 20,466 | 23,811 | 22,304 | 0 | 21,427 | 38,767 | 20,488 | 262,262 | |
| 1700--49 | files | 0 | 10 | 11 | 10 | 11 | 10 | 14 | 10 | 0 | 10 | 77 | 10 | 173 |
| words | 0 | 27,862 | 44,057 | 21,511 | 23,265 | 21,315 | 22,066 | 21,612 | 0 | 20,812 | 33,896 | 20,495 | 256,891 | |
| 1750--99 | files | 10 | 10 | 10 | 10 | 10 | 10 | 20 | 10 | 0 | 10 | 70 | 11 | 181 |
| words | 25,386 | 27,484 | 45,198 | 21,752 | 21,284 | 20,367 | 21,002 | 23,172 | 0 | 20,599 | 29,589 | 23,043 | 278,876 | |
| 1800--49 | files | 10 | 10 | 10 | 10 | 11 | 10 | 10 | 10 | 0 | 10 | 25 | 10 | 126 |
| words | 30,804 | 31,211 | 45,107 | 21,777 | 23,249 | 20,531 | 20,286 | 22,951 | 0 | 21,015 | 12,671 | 20,883 | 270,485 | |
| 1850--99 | files | 10 | 10 | 10 | 10 | 10 | 10 | 10 | 10 | 0 | 10 | 26 | 10 | 126 |
| words | 30,684 | 34,856 | 43,427 | 21,322 | 21,243 | 20,757 | 22,265 | 23,072 | 0 | 21,810 | 10,819 | 21,789 | 272,044 | |
| 1900--49 | files | 10 | 11 | 10 | 10 | 10 | 10 | 10 | 10 | 0 | 10 | 29 | 10 | 130 |
| words | 26,717 | 31,391 | 45,408 | 21,123 | 22,208 | 21,160 | 20,213 | 21,977 | 0 | 21,664 | 12,529 | 22,424 | 266,814 | |
| 1950--99 | files | 10 | 11 | 10 | 10 | 10 | 10 | 13 | 10 | 0 | 10 | 28 | 10 | 132 |
| words | 23,437 | 32,200 | 45,109 | 21,093 | 22,723 | 20,721 | 20,994 | 22,935 | 0 | 21,385 | 11,361 | 22,060 | 264,018 | |
| TOTAL | files | 50 | 82 | 72 | 70 | 72 | 80 | 98 | 70 | 10 | 70 | 330 | 71 | 1,075 |
| words | 137,028 | 247,674 | 309,973 | 150,396 | 155,158 | 166,343 | 150,637 | 158,023 | 32,741 | 148,712 | 149,632 | 151,182 | 1,957,499 | |
| AMERICAN | a | d | f | h | j | l | m | n | p | s | x | y | TOTAL | |
| 1750--99 | files | 3 | 10 | 10 | 10 | 10 | 12 | 9 | 10 | 0 | 10 | 58 | 10 | 152 |
| words | 9,214 | 29,980 | 38,980 | 21,271 | 21,896 | 41,177 | 23,541 | 22,265 | 0 | 20,668 | 27,860 | 21,315 | 278,167 | |
| 1800--49 | files | 1 | 10 | 10 | 0 | 10 | 12 | 0 | 10 | 0 | 10 | 10 | 10 | 83 |
| words | 2,822 | 40,568 | 44,676 | 0 | 21,476 | 33,409 | 0 | 37,107 | 0 | 20,904 | 20,739 | 20,695 | 242,396 | |
| 1850--99 | files | 8 | 10 | 11 | 10 | 10 | 10 | 10 | 10 | 0 | 10 | 28 | 11 | 128 |
| words | 24,480 | 32,721 | 44,394 | 21,056 | 22,436 | 28,506 | 20,547 | 21,994 | 0 | 21,311 | 11,361 | 23,419 | 272,225 | |
| 1900--49 | files | 10 | 10 | 10 | 0 | 10 | 11 | 0 | 15 | 0 | 10 | 52 | 10 | 138 |
| words | 30,460 | 52,514 | 53,430 | 0 | 21,661 | 21,607 | 0 | 22,802 | 0 | 20,984 | 25,021 | 20,731 | 269,210 | |
| 1950--99 | files | 10 | 10 | 10 | 10 | 10 | 12 | 10 | 10 | 0 | 12 | 30 | 10 | 134 |
| words | 29,563 | 31,037 | 44,382 | 21,051 | 22,109 | 25,517 | 22,617 | 23,069 | 0 | 25,623 | 11,961 | 21,654 | 278,583 | |
| TOTAL | files | 32 | 50 | 51 | 30 | 50 | 57 | 29 | 55 | 0 | 52 | 178 | 51 | 635 |
| words | 96,539 | 186,820 | 225,862 | 63,378 | 109,578 | 150,216 | 66,705 | 127,237 | 0 | 109,490 | 96,942 | 107,814 | 1,340,581 | |
Powered by WinChalow1.0rc4 based on chalow