hellogѸ˥֥     ChangeLog ǿ    

representativeness - hellogѸ˥֥

ǽ: 2026-07-15 01:27

2020-03-07 Sat

#3967. ѥѤ (3) [corpus][methodology][representativeness]

ɸˤĤƤϡʲεޤ͡ʵ˼夲Ƥ

#307. ѥѤ ([2010-02-28-1])
#367. ѥѤ (2) ([2010-04-29-1])
#428. The Brown family of corpora Ѿա ([2010-06-29-1])
#1280. ѥɽ ([2012-10-28-1])
#2584. ˱Ѹ쥳ѥɽ ([2016-05-24-1])
#2779. ѥϱѸ˸˻Ȥ뤱ɤ ([2016-12-05-1])

ѥѤѸʻˡ˸ϤޤޤˤʤäƤƤꡤسǤ뤵褦ˤʤä餳ѤˤäǧƤȤǤ롥ݤϤ褽֤Ȥʤ뤬ϱѸγ路 Fischer et al. (14) ꡤ4Ŧ褦

(i) there can be tension between what is easily retrieved through corpus searches and what is thought to be linguistically most significant; a historical syntactic case in point involves patterns of co-reference of noun phrases . . . ; these have been largely neglected because they involve information status, which is currently not part of any standard annotation scheme;

(ii) when a data search yields large numbers of hits, there may be a temptation to interpret corpus results merely as numbers, which is a severely reductive approach; in cases of grammaticalization, for example, changes in frequency may act as tell-tale signs . . . , but an exclusive quantitative focus will mean that one is ignoring the changes in meaning and context that form the core of the process;

(iii) the substantial amounts of data that can be collected from a corpus can also blind researchers to the dangers of making generalizations about the language as a whole on the basis of a partial view of it; this is a particularly relevant problem for diachronic research, because we only have very incomplete evidence for the state of the language in any historical period . . . ;

(iv) trying to achieve greater representativness by collecting and comparing data from various corpora can also be tricky: principles guiding text inclusion vary widely, there is little standardization in user interfaces, and they can require a significant time investment to learn to operate.


4θդĶСΤ褦ˤʤ롥

(i) ѥǿԤ䤹꤬Ūˤɬ̣ΤǤϤʤ⤷ʤդ٤
(ii) ŪʴŻ뤹븦ˤΩŪʴᤴƤޤ
(iii) ʥѥǤäȤƤ⡤ representative Ǥ櫓ǤϤʤʤ˸ؤˤ "bad-data problem"
(iv) ѥԻԤ䥤󥿡եԤΰտޤĤǡˡ򿴤ƽϤ٤

Fischer, Olga, Hendrik De Smet, and Wim van der Wurff. A Brief History of English Syntax. Cambridge: CUP, 2017.

Referrer (Inside): [2025-06-20-1] [2022-10-12-1]

[ | ѥڡ ]

2018-07-21 Sat

#3372. űѸѸλˤĤƿΥ [oe][me][philology][manuscript][statistics][representativeness][methodology][evidence]

#1264. ˸ؤθ³ȡιؤƻ ([2012-10-12-1])#2865. Ĥ䤹ڵ򡤾ä䤹ڵ򡽡ؤΥҥȡ ([2017-03-01-1]) Ǽ夲Ƥ˸ؤˤϻθ³ȤȤ⤷꤬롥̤Ȥ˾ۤɤΤΤĤäƤƤʤΤ¤Ǥ롥 (29) ϡ˸ء٤ΤʤΡָűѸαŪŪפȤˤƼΤ褦˽Ҥ٤Ƥ롥

˸ؤǤϸ¸ǽפǤ뤳ȤϤޤǤʤθ򸦵оݤȤΤǤСʸϿ˲äơüԤθľѡʤʤɤθ쿴ŪޤƼ¤˭٤ʻȤΤǤ뤬˸ؤǤϤϴñˤʤʤŤλȤȤ¿ʬʤΤǡμθ³Τ褦˻פ뤬ºݤϡʾ˸󤬤ΤǧʤФʤʤä˰ŤиŤۤɸΤꡤѸˤǤϡäOEθ³ˤĤƤϤ褯ǧǡʤƤʤƤϤʤʤ


Ūˤɤ줯餤󤬤Τָ뤿ˡűѸѸλˤĤƻϤƤս⤷Ƥ

űѸμܤ˴ޤޤ300ǡʸ2000Ǥ롥ʬŪˤϥ󥰤ˤ˲Ǥ롥̤ϥΥޥʹߤ200ǯ֤˽񤫤줿Ѹλ⾯ʤ(29)
Ȥ850ǯλǻĤäƤΤϡ4ĤΥƥȤ35ۤɤˡΧʸļʤɤûŪʸȾǤ롥(29)
űѸ9䤬ȥǽ񤫤줿Ǥ롥(31)
ɮ (authorial holograph) ѸǤ Ayenbite of Inwit 1340ǯˡ Hoccleve (1370?--1450?) νʪ15Ρեε²μˤʽ Paston Letters 䤽¾ƱνʽۤɤǤ롥(34)

Ѹˤˤ礭ʸ (philology) ʸɾ (textual criticism) Υץʬǽ׻뤵ʤǤ롥
Ϣơ#1264. ˸ؤθ³ȡιؤƻ ([2012-10-12-1])#2865. Ĥ䤹ڵ򡤾ä䤹ڵ򡽡ؤΥҥȡ ([2017-03-01-1])#1051. Ѹ˸оݤȤʤ (1) ([2012-03-13-1])#1052. Ѹ˸оݤȤʤ (2) ([2012-03-14-1]) ⻲ȡ

2ϡѸ˳ѡ ԡˡ˸ءīоȸإ꡼ȯŸԡ3īҽŹ2018ǯ22--46ǡ

[ | ѥڡ ]

2017-03-01 Wed

#2865. Ĥ䤹ڵ򡤾ä䤹ڵ򡽡ؤΥҥ [philology][writing][manuscript][representativeness][textual_transmission][evidence]

ˤ򸦵椹ۤͣˡϡ¸˰͵򤹤뤳ȤǤ롥Ȥ¸ϼŪˤŪˤФ꤬ꡤѥؤѸǤ "representative" Ǥ "balanced" Ǥʤ
ޤʪŪʾ郎롥ߤ뤿ˤϡĹ֤Ѥ̺ƻǵƤ뤳ȤɬפǤʡ#2457. ̺Ƚƻ (2) ([2016-01-18-1]) 򻲾ȡˡˡǥδ顤äաŪʸոˤϡ񤭻Ĥơߤޤǽ㤤Ȥ餫񤭼δϡʸˤϤäѤɤ߽ǽϤΤ륨꡼ؤθո乥ߤ꤬ȿǤ졤ʳؤθưϿ˻Ĥ뤳ȤϤۤȤɤʤˡƤȤ顤㤨ȿŪʽʪϡˤǽ⤤
¸ʸϡ͡ʱ̿򤹤ȴĤäƤȤ̣ˤΡֶפθƤΤϳΤĤ䤹ΤˤϤĤξ郎Ȥ嵭εǰƬ֤ȡΡɬפθƤȤ롥˸ؤʸؤˤơɤΤ褦ʡ־פɬפΤƤȤϽפ
ΤδŪʥҥȤȤơɤΤ褦ʲФפߤޤĤ䤹ͻ벽 (taphonomy) Ȥʬθ򻲹ͤˤ㤨СಽФλĤ䤹ĤˤϡˤäƷޤȹͤåɤϡֿಽеϿζкߡפꤹ (73--75) ǡΤ褦˽Ҥ٤Ƥ롥

ǯ֤⡤ؼԤϡ700?600ǯʹߤ¸οΤͳ褹벽Ф򽸤Ƥο¿褦˻פ뤫⤷ʤʬϸߤ˶ᤤǯΤΤǯŪкߤΤۤˤ⡤ಽФˤϤޤޤФ꤬롥Τ褦Ф꤬뵡餫ˤȤϡֲءפȤФ롥䲼ܹ뤤ϻͻ礭ʹϤ褯Ĥ뤬ǹϾסؤιʤɤ̩ưפ»ΤǻĤˤĤޤꡤλĤ䤹礭ȴ椵㤹롥ǹΤ褦ʷڤϱˤή졤Ф˱Ф졤ʤˤιȰ˲Фˤʤ롥ŤƬܹϹή졤δδ֤˰ääơۤΦ緿ưʪιȰ˲Фˤʤ롥
ФλĤ򺸱⤦ĤװϡῩưʪΤΤɤʬ򹥤फǤ롥ҥ祦ϥμ­ΤʤΤǡ⤷ΤǤῩưʪƱäΤʤ顤ಽФμ­ȯ뤳ȤʤϤǤ롥ºݡ­βФϤޤ껺ФƤ餺ΤοʲˤĤƤϤ褯狼äƤ뤬­οʲϤ褯狼äƤʤΤ礭⡤ФĤ뤫ɤ˱ƶ롤Τ礭ʼϲФȤƻĤ䤹ƱǤΤ礭ʸΤϻĤ䤹󡤤Τ褦ФϿಽФ⳺롥
ĶǤϡۤ١Фˤʤ䤹ȯ䤹äơǯ䤢ϰͳ褹벽Ф¿Ȥäơǯ夢뤤ϰ¿θΤǤȤϸ¤ʤǯϰξۤ٤ƲвŬƤǽΤƱͤˡǯϰ褫ಽФĤʤƤ⡤˿बǤʤäȤˤϤʤʤ־ڵΤʤΤϡ¸ߤʤڵǤϤʤפȤʸ⤢롥äơΤμϡǸŤβФȯǯǿβФȯǯޤ¸ƤȤˤʤ롥ĤޤꡤмεǤǯϼºݤ٤ƾ˹ˤʤ롥
Ʊͤ¤ϲȯפŪʬۤˤƤϤޤ롥ϡФȯפ깭ϰϤ¸ƤϤޤδĶϸߤȤϰäƤϤߤǤϸĶˤϽߤ䤹ä⤷ʤդ⤢롥ˡФȤ¸ƤĶϷ褷¿ϤʤھǤϡä˻ĤʤȤ˿ӴĶǤϡ٤⤯ھʤΤǡФϻĤʤȹͤƤǶǤϡɬ⤽ǤϤʤȤ餫ˤʤäȤϤ͸ųؼԤдȿ͹ȯȴäƤ⡤Ƥ͹ϤƤޤд路ȯǤʤΤ


ϡνФѸȤʤäʸؾȤϢ˸ˤˤĤƤϡ#1051. Ѹ˸оݤȤʤ (1) ([2012-03-13-1])#1052. Ѹ˸оݤȤʤ (2) ([2012-03-14-1]) 򻲾Ȥ줿ϢơʸȯˤĤƤιͻ⻲ͤˤʤʡ#1834. ʸǯɽ ([2014-05-05-1])#2389. ʸηϤεȯã (1)([2015-11-11-1])#2457. ̺Ƚƻ (2) ([2016-01-18-1]) 򻲾ȡˡ

СʡɡåɡˡϾ ͪˡˡؿοʲȻǤˤõ١ǡ2014ǯ

[ | ѥڡ ]

2016-12-05 Mon

#2779. ѥϱѸ˸˻Ȥ뤱ɤ [hel_education][corpus][methodology][philology][representativeness]

˸ؤȥѥѤȤ褤ȸ롥㤨бѸˤǹͤСŪʥƥȤͭ¤Ǥ뤫顤٤ƤŻŪ˳ǼСʸ̤ŪĴ뤳ȤǤ롥αѸ˵Ҹ⡤ŻŪǤʤäΥƥȷĴоݤȤƤǤϡƱ֥ѥ׸äȤ⤢롥ߤǤϡޤޤ¿θԤŻҥѥȤδϢġѤƸ椷Ƥ뤷Ĭήϴְ㤤ʤȯŸƤ
դ٤Ȥ⤢롥Ѹ쥳ѥؤ路ɥ (197) Τ褦˽Ҥ٤Ƥ롥

1990ǯʹߡ˸ؼԤˤȤäƥѥ˽פʥġȤʤޤȤʸʤФʤʤä⡤ʤΨ褯ܤõˤʤޤȤϤ˸ؤǤϡΥѥȤäʾˡʸΩ֤ꡤɬפޤʤŤǤϡΡָʸפɬ⸶ܤǤϤʤʣ̤Ǥ뤫⤷줺ˤϤʣ̤⡤ʣμ̻Ȥ˼DzԤƻĤΤΰĤˤʤǽޤˡ顤μŪʥѥؤϡ󡤻ɴǯθθȤۤˤޤŤαѸ쥳ѥ򤷤ʤСѸˤμȤȤδϢǹԤΤ褤Ǥ礦


Ѥ˽Ҥ٤Ƥ̤ꡤȤ櫓ŤθоݤȤˤϡդפ롥θ򰷤Τ褦ΤޤޤǤȡˤ֤Ĥ뤫㤨СѸˤǤСѸǤֻꤷƤʤñ򸡺褦ȤƤ⡤ɸŪֻʤΤǡͤ͡ʰ֤ǸƤߤɬפ롥ޤʸˡŪʥդʤƤˤ⡤ѸȸűѸ졦ѸȤǤʸˡƤۤʤäƤ뤷޸μ㤷ƤΤǡˤäƻ˸ŤʸˡμɬܤȤʤ롥פˡ˸Ť˽ϤƤ뤳ȤΤǤ롥
ˡѤǸڤƤ̤ꡤָʸפοԲķǤˤ⤫餺ѥɽŪѤ뤳Ȥ˴ƤȡƤοߤ˰ռʤʤ꤬Сʸɾ (textual criticism) ʸŪʥץ±ˤʤäƤޤ롥󡤥ѥΤȤ櫓ǤϤʤ椹ԤΤȤռŪ˼ФƤФ褤ȤȤǤϤΤñǤϤʤ
⤦1ġŻҥѥʤ餺Ȥˤʤ뤬ѥɽ (representativeness) ꤬ˤ롥ŤξˤϡߤƥȤ϶ĤäƤȤפΤʤ𤬤ꡤɽϤȤ˿
Ťθ򥳡ѥĴȤȤϡŪˤŪˤ⡤츫ۤɤ䤹ʤȤȤ򤷤ƤϢơʲε⻲ȡ

#568. ѥȱѸ쥳ѥ ([2010-11-16-1])
#363. Ѹ쥳ѥȯŸ3 ([2010-04-25-1])
#368. ѥϸβǽ򹭤 ([2010-04-30-1])
#1165. ѹǥѥ椬ˤʤäطʡ ([2012-07-05-1])
#307. ѥѤ ([2010-02-28-1])
#367. ѥѤ (2) ([2010-04-29-1])
#1280. ѥɽ ([2012-10-28-1])
#2584. ˱Ѹ쥳ѥɽ ([2016-05-24-1])

ϡ󥹡ɥȡˡ 翹 ʸҡ ޤߡ ɹˡرѸ쥳ѥѤ츦١罤ۡ2016ǯ

[ | ѥڡ ]

2016-06-07 Tue

#2598. ťΥɸαƶϤŤõ븦ˤαդ٤Ѹ쥳ѥ [old_norse][loan_word][me_dialect][representativeness][geography][lexical_diffusion][lexicology][methodology][laeme][corpus]

#1917. numb ([2014-07-27-1]) εǡѸˤ nimen ȸťΥɸѸ taken ζˤĤĴ Rynell θ˿줿̤˸ťΥɸѸ줬Ѹ椤ˤƱѸ˿ƩƤäݤˤϡδϰδθ롥ΤȤʤ顤οƩˤϤ٤λ֤ΤǡΤۤƩٹ礤ϸȤʤޤťΥɸαƶ the Danelaw ȸƤФ륤󥰥ˤƺǤ⶯Ǥꡤ󥰥ؤϡξ׷⤬֤ޤʤŤƤäȹͤΤǤ롥
Τ褦˸ťΥɸθŪƶζˤĤƤϡϰδ֤̩ܤߴطꡤʬۤΤǤȤ롥ºݤˡ#818. 󥰥ɤ˻ĤťΥɸ̾ ([2011-07-24-1]) #1937. Ϣ -son ˤΤϸťΥɸͳ ([2014-08-16-1]) ˼ʬۿޤϡΤʬۤ򼨤űѸȸťΥɸѸ줬礹륱Ǥϡ̤˾嵭ʬۤǧ뤳Ȥ¿褦Rynell (359) ۩"The Scn words so far dealt with have this in common that they prevail in the East Midlands, the North, and the North West Midlands, or in one or two of these districts, while their native synonyms hold the field in the South West Midlands and the South."
ϰ츫ۤñǤϤʤȤˤαդɬפ롥Rynell (359--60) Ͼʸ³ơΤ褦â񤭤դäƤ롥

This is obviously not tantamount to saying that the native words are wanting in the former parts of the country and, inversely, that the Scn words are all absent from the latter. Instead, the native words are by no means infrequent in the East Midlands, the North, and the North West Midlands, or at least in parts of these districts, and not a few Scn loan-words turn up in the South West Midlands and the South, particularly near the East Midland border in Essex, once the southernmost country of the Danelaw. Moreover, some Scn words seem to have been more generally accepted down there at a surprisingly early stage, in some cases even at the expense of their native equivalents.


äդ٤ϡ¸ѸƥȤʬۤФäƤǤ롥򤫤СѸ쥳ѥϰ˴ؤɽ (representativeness) 礤ƤȤRynell (358) ˤС

A survey of the entire material above collected, which suffers from the weakness that the texts from the North and the North (and Central) West Midlands are all comparatively late and those from the South West Midlands nearly all early, while the East Midland and Southern texts, particularly the former, represent various periods, shows that in a number of cases the Scn words do prevail in the East Midlands, the North, and the North (and sometimes Central) West Midlands and the South, exclusive of Chaucer's London . . . .


ťΥɸθŪƶϡѸᤤǡ٤ˤǴѻ롤ȤȤϳȤƽҤ٤뤳ȤϤǤΤΡ줬Ѹ쥳ѥλʬۤȸ˰פƤ¤ƨƤϤʤʤĤޤꡤ嵭γŪʬۤϡޤ޸¸ƥȤλ֡ŪʬۤʿԤƤ뤿ˡȤˤ˶ĴƤ뤫⤷ʤΤ䤹Τޤޤ䤹ʤꡤˤΤ줿ޤޤˤ빽¤Ū꤬ˤ롥
ϡťΥɸθŪƶˤȤɤޤ餺Ѹ̡ŤѲ̤ѻݤˤͿǤ (see #941. ѸθѲϤʤ̤ŤΤ ([2011-11-24-1])#1843. conservative radicalism ([2014-05-14-1]))
ϢơѸ쥳ѥ A Linguistic Atlas of Early Middle English (LAEME) ɽˤĤơ#1262. The LAEME Corpus ɽ (1) ([2012-10-10-1])#1263. The LAEME Corpus ɽ (2) ([2012-10-11-1]) ⻲ȡ

Rynell, Alarik. The Rivalry of Scandinavian and Native Synonyms in Middle English Especially taken and nimen. Lund: Håkan Ohlssons, 1948.

[ | ѥڡ ]

2016-05-24 Tue

#2584. ˱Ѹ쥳ѥɽ [representativeness][corpus][methodology][hc][register]

ѥɽ (representativeness) ѹ (balance) ˤĤƤϡ#1280. ѥɽ ([2012-10-28-1]) ¾εǰäƤŤαѸΥѥ򰷤ˤϡѸ쥳ѥ˴ؤ꤬ΤޤƤϤޤΤΤȤʤ顤˾褻Ƥ˺꤬¿ΩϤ롥
ޤ˱Ѹ쥳ѥȤȤơѥθʤ콸ΥƥȽΤΤΤˤζˤ긽¸ƤΤ˸¤Ȥ󤬤롥ʸܡܡʤɤ˵ƸߤޤĤꡤ¸ƤΤ٤ƤǤ롥ޤλ¸ߤ뤳ȤϤ狼äƤƤ⡤Ū˥Ǥ뤫ɤǤ롥Ūˤϡ뤤Żҷ֤ǽǤƤ뤫ɤˤäƤλƥȤ콸Ȥʤꡤ褯ԻԤˤäΰѥʼȤŻҷ֡ˤؤԻ뤳Ȥˤʤ롥Ω˱Ѹ쥳ѥϡ򤫤äơ褦䤯˽ФΤǤꡤλŪɽãƤ븫ߤϡǰʤ
ޤ˱ѸȤҤȤ˸äƤ⡤ºݤˤϸѸƱͤ͡ lects registers ʬ졤ζʬ˱ƥѥԻ륱¿Τˡ̣ѥѥȸƤǤ褤 Helsinki Corpus Τ褦̻ѥ䡤òƤ뤬ۤʤޤ Penn Parsed Corpora of Historical English ⤢뤷ȤˤäƤ̻ѥȤƤѤǤ OED ΰʸʤɤ롥̾ϡԻŪ֤˱ơ꾮ϰϤΥƥȤ˹ʤäԻ륳ѥ¿űѸ쥳ѥѸ쥳ѥʤɻˤäƶڤ뤳 (chronolects) ⤢ ꥹѸ䥢ꥫѸʤɤ (dialects) ξ⤢뤷Chaucer Shakespeare ʤκ (idiolects) ξ⤢Ҳ (sociolects) ̤Ȥ⤢뤷Ѱ (registers) ˱ƥѥԻȤȤ⤢롥ѰȤäƤ⡤äξʥˡΡäդ񤭸դˡʷˤʤɤ˱ơ̶ʬ뤳ȤǤ롥ʬ򤢤ޤ٤ƤޤȡҤΤ褦˸¸ƥȤ̤ͭ¤ǤꡤƤ˾ʤäʬۤФäƤ櫓顤ɽѹդݤĤȤʤΤȺȤʤ롥
(chronolects) μ濴ˤƶǯŪ絬Ϥ˱Ѹ쥳ѥԻ򳵴ѤƤߤȡƻαѸμʸˡϿޤΤ褦ʻͻԻȴϢŤԻ줿ΤĤ뤳Ȥ狼롥ϡƻδ𼴥ѥȤư֤ŤȤäƤ褤⤷ʤ㤨СDictionary of Old English Corpus (DOEC), A Linguistic Atlas of Early Middle English (LAEME), Middle English Grammar project (MEG) Ǥ롥ˤĤƤϥѥȤϥƥȡǡ١Ȥ٤ EEBO (Early English Books Online) ѲǽȤʤäƤƤ뤷ꥫѸˤĤƤ Corpus of Historical American English (COHA) λߤ⤢롥
˵󤲤ɽŪ̾ʤΤΤۤ͡ڤ˱Ѹ쥳ѥԻδ褬³ȸƤ롥ˤĤƤϡ#506. CoRD --- Ѹ˥ѥξ󥻥󥿡 ([2010-09-15-1]) ǾҲ𤷤Helsinki ؤ VARIENG ( Research Unit for Variation, Contacts and Change in English ) ץȤ CoRD ( Corpus Resource Database ) 򻲾Ȥ줿
ϡŪˤСҤΤ褦ˡ˱Ѹ쥳ѥɽŪ˲褹ΤϺǤϤ뤬Ǹ̥ѥԻϳ趷褷ƤꡤΤʤǿⴶϤ롥ҤȤĤϥեŪ¦̡Ѽ¦ΥѥФ٤˴ؤ⤢褦˻פ롥Żҥѥλ夬褹ˤ⡤˱ѸθԤϡ֤פȤʤΤˤ֥ѥפԻѡʬϤ˹ԤʤäƤΤǤ롥⡤䤿ռƤۤɤǤϤʤäΤΡ٤ϥѥɽѹդȤäͤƤΤǤꡤ˻ʤȤƤ⡤ͭפʸ³θѤƤŻҥѥȤʤäɽѹդ꤬ΩäƼ夲褦ˤʤäΤˤ꼫Τڤǡͤ³ɬפϤΤΡ餫ˤ츽ݤΤΤ˾ơѥ˻Ȥʤʤ餽θ³Ƥ椯ȤפʤΤǤϤʤ

[ | ѥڡ ]

2016-03-22 Tue

#2521. Ѹ113 "such" ΰֻ [spelling][eme][laeme][corpus][scribe][me_dialect][representativeness]

ε#2520. Ѹ134 "such" ΰֻ ([2016-03-21-1]) ³ϽѸ쥳ѥ LAEME "such" ΰֻФƤߤ (see #1262. The LAEME Corpus ɽ (1) ([2012-10-10-1])) θϡѸǤϷƻ졤졤³ȤѤ졤ƻξˤ϶ޤ⤹ΤǡΤȤ͡ʷ֤롥ե٥åȽ˰褦ʤäοͤʸڤ١ˡ

hsƿucche (1), schilke (1), schuc (3), scli (1), scuche (1), sec (1), secc (1), secche (1), sech (2), seche (1), selk (1), selke (1), shuc (1), shuch (1), siche (1), silc (1), silk (3), sli (1), slic (5), sliik (1), slik (3), slike (1), slk (1), sly (1), soch (5), soche (1), solchere (1), suc (2), sucche (2), such (51), suche (1), suecche (1), suech (1), sueche (1), sueh (1), sug (1), suic (1), suicchne (1), suich (12), suiche (3), suilc (14), suilce (1), suilch (1), suilk (1), suilke (2), sulch (1), sulche (1), sulk (1), sulke (1), suuche (1), suweche (1), suwilk (1), suyc (1), suych (4), suyche (1), svich (2), sƿche (1), sƿic (3), sƿicche (1), sƿich (1), sƿiche (14), sƿichne (1), sƿilc (30), sƿilch (22), sƿilche (1), sƿilcne (1), sƿilk (14), sƿillc (10), sƿillke (2), sƿi~lch (1), sƿlche (1), sƿuc (4), sƿucch (1), sƿucche (4), sƿucches (1), sƿuch (1), sƿuche (1), sƿuchne (1), sƿuilc (1), sƿulc (8), sƿulce (1), sƿulche (9), swch (1), swecche (1), swech (1), sweche (2), swich (5), swiche (1), swics (1), swil (1), swilc (5), swilce (1), swilk (2), swilke (2), swilkee (2), swlc (1), swlch (1), swlche (1), swlchere (1), swlcne (1), swuche (2), swuh (1), swulcere (1), swulch (3), swulchen (1), swulchere (1), swulke (1), swulne (1), zuich (10), zuiche (14), zuichen (3), zuych (10), zuyche (2)


ʸȾʸζ̤ϤĤˡ113ֻʸڤ롥Τʤ٤ˤƥȥå5ֻȴФȡsuch, sƿilc, sƿilch, sƿilk, sƿiche Ȥʤꡤ5369ĤΤ131 (35.5%) 롥
θѸ줫134ȹ碌ʣֻ򸺻ȡѸΤȤ247ΰֻ뤳Ȥˤʤ롥ѤϿޤ䥳ѥɬŪǤϤʤΤǡϹʿͤȻפ롥㤨СMED swich (adj.) ˷ǤƤֻäСϤ⤦

Referrer (Inside): [2018-08-16-1]

[ | ѥڡ ]

2014-01-30 Thu

#1739. AmE-BrE Diachronic Frequency Comparer [corpus][ame_bre][web_service][cgi][frequency][representativeness]

#1730. AmE-BrE 2006 Frequency Comparer ([2014-01-21-1]) ǡ2006ǯν񤭸եƥȤԻƳѼ拾ѥҲ𤷡˴ŤӥġΥġʤ鵤ŤΤƱˡԻ졤ϤƱ100٤ the Brown family of corpora ʡ#428. The Brown family of corpora Ѿա ([2010-06-29-1])ˤϢȤСľ50ǯ֤ۤɤ̻ŪʱƴӤưפ˲ǽȤʤ롥
ǡεǾҲ𤷤 Professor Paul Baker - Linguistics and English Language at Lancaster University ˤ AmE06 BrE06 ˲äơ񤭸եꥫѸɽ Brown (1961), Frown (1992)񤭸եꥹѸɽ LOB (1961), FLOB (1991) ɽФ碌ƥǡ١ѤλϡAmE-BrE 2006 Frequency Comparer ȤۤƱʤΤǡμ ([2014-01-21-1]) 򻲾Ȥ줿ϤɽǤϡθиƥȤοٽ̤ϾʤƤꡤ100٤ɽˤȤɤƤΤǡAmE06 BE06 ˤĤԤξɬפʾˤϡAmE-BrE 2006 Frequency Comparer ɤ

    
Sort: by Brown freq by LOB freq by Frown freq by FLOB freq by AmE06 freq by BE06 freq alphabetically nothing (non-regex mode only)

㤨С^movies?$ ϤƤߤȡŪ˥ꥫѸŪȤƤθʬۤ50ǯۤɤδ֤ˡꥹѸˤ⿻ƩƤƤͻҤ狼롥
ƺ̻ŪѲĴΤǤСñǤϤʤĤĵϤʡ#607. Google Books Ngram Viewer ([2010-12-25-1]) ΤۤؤΥġϡthe Brown family of corpora ١ˤƤ뤬椨ˡ(1) ѹդӲǽǤꡤ(2) פ狼äƤʺƸǽݤƤˤȤ뤳ȤϻŦƤ˾ޤΤϡǤ٤ʥѥȡ緿ǷŤߤˤ륳ѥȤϢȤ뤳Ȥ

[ | ѥڡ ]

2014-01-07 Tue

#1716. shew show (3) [spelling][corpus][clmet][representativeness]

#1415. shew show (1) ([2013-03-12-1]) ȡ#1416. shew show (2) ([2013-03-13-1]) ǰä򡤡#1637. CLMET3.0 between betwixt ʬۤĴ ([2013-10-20-1]) ǾҲ𤷤 The Corpus of Late Modern English Texts, version 3.0 (CLMET3.0) ˤˬŪˤϡƱդѥ "\bshow(s|n|ed|ing)?_VB" "\bshew(s|n|ed|ing)?_VB" Ǹơ3ʬȤٿ٤ʲη̤Ф


shew show
1710--17803351,54510,480,431
1780--18501593,10011,285,587
1850--1920925,11812,620,207


#1416. shew show (2) ([2013-03-13-1]) Ѥ PPCMBE (Penn Parsed Corpus of Modern British English) 100Υѥ CLMET3.0 3,400ε祳ѥǤ롥ۤƱ򥫥СƤΤӤˤԹ礬褤Ʊͤ˺ show shew ¤֤ƤͻҤ뤬礭ۤʤΤϡ1710--1780ǯ1ˤƤǤ show Ū˾äƤ뤳ȤǤ롥򿮤ʤСѸޤǤˡǤ show ϾԤ褷ƤȤȤˤʤ롥PPCMBE Ǥ shew ϸѸˡͥƱפȿܤCLMET3.0 Ǥϡ餫äפȿܤƤ롥2ˤ錄̻ŪʻξѥȤ绨Ĥˤϻ褦ʷ򼨤ȤϤΤΡ18ζŪʬۤˤĤƤξѥμͤκ礭褦˻פ롥ˤϡ#1280. ѥɽ ([2012-10-28-1]) Ȥ꤬ؤäƤǤꡤŤʲ᤬뤳Ȥˤʤ
ʤ1Ĥʸ̮ shew show ȤѤƤ붽̣⤤Ĥä3Τߵ󤲤褦

Why, you have shewn your wit upon the subject, and I mean to show your courage;
Mr. Wright, as well as Nadin, professed they were perfectly satisfied of this, and appeared to shew to me all the polite attention that they were capable of showing.
Assuredly I did not show him the face which I shewed Folderico.

Referrer (Inside): [2019-10-15-1] [2014-04-07-1]

[ | ѥڡ ]

2012-10-28 Sun

#1280. ѥɽ [corpus][representativeness][variety][idiolect][methodology]

ѥˤȤäɽ (representativeness) ̿Ǥ뤳Ȥϡѥ ([2010-11-16-1]) 餫Ǥ뤷ε#1279. BNC ζߤȼߡ ([2012-10-28-1]) ǾҲ𤷤 Leech Ȥ櫓ĥƤǤ롥McEnery et al. (13) ϡɽˤĤơLeech 򻲹ͤˤʤ "a corpus is thought to be representative of the language variety it is supposed to represent if the findings based on its contents can be generalized to the said language variety" ȽҤ٤Ƥ롥
ɽŪ˹ͤƤߤ褦㤨 BNC åȤȤ褦ʡ奤ꥹѸȤŪѼϿ륳ѥ (general corpus) ɽϤɤΤ褦ˤΤ񤷤äդȽ񤭸դγͤȡ줾50%Ĥ˳꿶뤳Ȥϡ奤ꥹѸɽ«ƤLeech ɽǤ "impressionistic" Ȥʤ餶ʤνִ֤˹ԤʤƤ븽奤ꥹѸΰŪʬäդˤƤǤϤʤ⤷ȤСäեѥγ㤨80%ۤɤꤹۤɽݤǤΤǤϤʤΤȤʤ븽奤ꥹѸľܤĤळȤǤʤʾ塤ɽεϹԤͤޤäƤޤ
ѥä˰̥ѥˤɽȤˡ balance sampling Ȥ2Ĥγǰʬƹͤ뤳Ȥ롥McEnery et al. (13) Ǥϡ"the representativeness of most corpora is to a great extent determined by two factors: the range of genres included in a corpus (i.e. balance . . .) and how the text chunks for each genre are selected (i.e. sampling . . .)" Ƥ롥
balance ȤϡBNC ѸǤȤ domain genre Ȥʬ˴ؤΤǤ롥㤨С奤ꥹѸΥѥɸ֤ʤ⡤ꥹοʹαѸ򽸤᤿ѥϡrepresentativeness 񤬤롥奤ꥹѸˤϽ񤭸դǤʤäդ⤢뤷ԤˤĤƤϿʹѸǤʤʸرѸ⤢Żҥ᡼Ѹ⤢뤷㤤ʪѸ⤢СѸ⤢롥Τ domain genre θ줿Ȼפ̤ƤĤ text domain ΤʹѸ˸¤äƤ⡤֥ɤ⤢й⤢롥1ĤοʹǤ⡤Ҳ̡ݡ̡ʤɤ̤ɬפϤʤΤҲ̤Ǥй⵭ȹݵζ̤ϤɤŪˤϤɤޤǤʬ롥äդǤƱͤ˺ʬ䤷ʤƤСĿ͸ (idiolect) ˸Ŀ͸ˤ register ̤θ졤ʤɤΥȥؤȽ夷ƤޤºݤΥѥϡQŪʥ٥Ŷ뤳Ȥˤʤ뤬־QŪפ "impressionistic" ϤۤƱ
sampling Ȥɽ뤿μˡǤ롥ΤθŪħƸ褦ˡ̤ˤƹθäʤ顤ѥ˳ domain ۤ뤿ȼǤ롥ˤϡsampling unit ȤƲꤹ뤫ŵŪˤϡܡʹʤɤʤȤƤñ̡ˡΤ褦ñ̤ꥹȲȤϰ (sampling frame) ɤޤǤꤹ뤫ǯؤθ䡤٥ȥ顼ܤؤθʤɡˡɸܼϴʥˤ뤫٤ηϲäǤΥˤ뤫ɤۤ뤫ʤɤΡŪŪ꤬ޤޤ롥
ɽ˴ؤ⤦1ĤγǰȤơclosure 뤤 saturation ȸƤФΤ⤢롥McEnery et al. (16) ˤС"Closure/saturation for a particular linguistic feature (e.g. size of lexicon) of a variety of language (e.g. computer manuals) means that the feature appears to be finite or is subject to very limited variation beyond a certain point." Ƥ롥ʿСʾ女ѥεϤ礭Ƥ⡤ùγѤʤȤϤãСΥѥ saturated Ǥȹͤ롥ɽλɸȤƤϡbalance saturation ΤۤƤȤŦ⤢뤬saturation ϼȤƸäǰƬˤꡤ¾θܤؤαѤϻߤƤʤΤǤ롥
ɽϡ女ѥ̿ǤȤϤäƤ⡤ԤȤ餤Ϥ롥ݤ뤿ʤˡʤ٤ƤΥѥԻԤΩϤƬˤѥϼԻƤ롥Ṳ̄ˤơҤԻȻѤ³Ƥ椭Υϥ򤿤٤ʳˤΤ⤷ʤ

McEnery, Tony, Richard Xiao, and Yukio Tono. Corpus-Based Language Studies: An Advanced Resource Book. London: Routledge, 2006.

[ | ѥڡ ]

2012-10-27 Sat

#1279. BNC ζߤȼ [bnc][corpus][representativeness]

108--114֤ˤ錄ꡤΩرѸ춵鸦ˤŤǡLancaster ̾ Geoffrey Leech θֱ񤬳줿ϡ2ܤ "The British National Corpus: Both a Triumph and a Failure" ꤹֱΤߤλääİ˹ԤäBNC ԼԤκäʤɡ⤷ää
̾ˤ "triumph" "failure" ˤĤơLeech Ϥ줾켡Τ褦ʹܤ󤷤Ƥ

A triumph:
It has been claimed that the BNC is the most widely used corpus in the world.
It was the first text corpus of its size to be made widely available.
It is available from a wide range of different sources.
It is widely regarded as a 'standard reference corpus' for the English language.
It has been licensed to over 1300 institutions throughout the world, over 1800 users have signed on for access to it through the BNCweb online interface, etc.

A failure:
It never reached 100 million words! (98,300,000)
The design criteria were never totally achieved.
It hardly ever contains complete texts.
The spoken materials are poorly transcribed.
The metadata are incomplete and can be erroneous.
The part-of-speech tagging contains many errors.
It is out of date! (dating from the late 20th century)


Leech θդüˤϡtriumph γ˼Ƥ褦ˡӤ΢դ줿ߤʤäƤΥѥԽˤĤơФ褫äФ褫äȤθȤ⤤ȿ¿󤲤ƤΤŪǤ롥BNC ΥդѤ줿ץ CLAWS4 ٤97%ۤɤHoffmann et al. 43 ˤȡ98--99%ˤȤΤϡ϶ä٤ȤȻפäƤѥϤ礭ΤǿѡȤΥ顼ȤϤäƤ300ˤΤܤȤ¤ϸȤƤäȤХѥˤĤƤϡѥΤ1ۤɤޤʤäȡǡ transcription μäȡѤǡեޥå TEI äȤФΥդˤɬŬڤǤʤäȡʤɤ󤲤Ƥ
ʤǤ⡤ʳ鸽ߤ˻ޤǰӤƤ³Ƥɽ (representativeness) ˤĤơBNC ǤϴṲ̄ʤäȤˡˤޤƤʳ顤ꤹ Text Domain ΥХ󥹤䥵˴ؤŤͤƤȤϤ褯ΤƤ롥1桼ȤƤϡ¤줿꥽ΤʤǡɽݤȤϰζȤɾƤ뤬Leech ˤȤäƤϡǤ¤ΤȤϤäȤȿ̤Ȥơ̤ۤʤäȤפ⶯褦Ʊˡ䤫ʸĴǤϤäBNC Ӥ¾Τ٤Ƥ絬ϥѥɽ򤵤ۤɽŻ뤷ƤʤȽƤ༫ȤҤ٤Ƥ褦ˡѥɽˤĤȼϤäƤ뤬ǽŪˤ "impressionistic" ȽǤȹͤƤ褦Ǥꡤ񤷤ˤޤƤˤ衤Leech ɽؤμǰζˡ٤ʥץեåʥꥺ򴶤
ʤ[2012-07-05-1]ε#1165. ѹǥѥ椬ˤʤäطʡפǿ줿̤ꡤǰʤBNC³ԤϤʤȤȤLeech Ƥ
礭ۤʤ뤬Ѹ쥳ѥ The LAEME Corpus ɽˤĤơ[2012-10-10-1], [2012-10-11-1]εǹͻΤǡȤ

Hoffmann, Sebastian, Stefan Evert, Nicholas Smith, David Lee, and Ylva Berglund Prytz. Corpus Linguistics with BNCweb : A Practical Guide. Frankfurt am Main: Peter Lang, 2008.

[ | ѥڡ ]

2012-10-12 Fri

#1264. ˸ؤθ³ȡιؤƻ [methodology][uniformitarian_principle][writing][history][sociolinguistics][laeme][corpus][representativeness][evidence]

[2012-10-10-1], [2012-10-11-1]εǡThe LAEME Corpus ɽˤĤƼꤢɾȤƤϡСƤȻȤߤɽ»ʤƤΤΡѤǤѸ쥳ѥȤƤηŪԤޤ줿絬ϤΥѥǤꡤʬդʧäǸ츦˳Ѥ٤ġǤ롥The LAEME Corpus β٤Ϥ󤢤뤷¾Υѥˤ䴰ܻؤ٤ȤϹͤ뤬Ū˸椹ݤɬŪˤĤޤȤ³θɾʤȥեǤ롥
˸ؤϡβξ֤ѻȤ򼫤˲ݤƤ롥򰷤Ȥˤϡߤ򰷤Ȥˤϸʤ³ĤޤȤMilroy (45) λŦ˸ظ2Ĥθ³ (limitations of historical inquiry) 򼨤

[P]ast states of language are attested in writing, rather than in speech . . . [W]ritten language tends to be message-oriented and is deprived of the social and situational contexts in which speech events occur.

[H]istorical data have been accidentally preserved and are therefore not equally representative of all aspects of the language of past states . . . . Some styles and varieties may therefore be over-represented in the data, while others are under-represented . . . . For some periods of time there may be a great deal of surviving information: for other periods there may be very little or none at all.


ۤ³ǤϤ뤬Ϥ뤤ϹˤǤŤϤϡˡǤʤƤ롥ΤʤǤ⡤Smith Ϥο (1) 񤭸դäդδط򿼤뤳ȡ(2) ̻ˤȳ̻ˤбܤ뤳ȡ(3) ߤθβؤαѤβǽõ뤳ȡνŦƤ롥
Ȥ櫓 (3) ˤĤƤϡǯҲؤˤѲ򤬵®˿ʤߡθβؤαѤˤʤ褦ˤʤäƤLabov ʸɸ "On the Use of the Present to Explain the Past" ˡľ٣ʪäƤ롥
ȴϢˡǤ uniformitarian_principle ưθ§ˤ̤˲Ф˱ѸʸDenison et al. ԽΤȤˡǯǤ줿ȤդäƤ

Milroy, James. Linguistic Variation and Change: On the Historical Sociolinguistics of English. Oxford: Blackwell, 1992.
Smith, Jeremy J. An Historical Study of English: Function, Form and Change. London: Routledge, 1996.
Labov, William. "On the Use of the Present to Explain the Past." Readings in Historical Phonology: Chapters in the Theory of Sound Change. Ed. Philip Baldi and Ronald N. Werth. Philadelphia: U of Pennsylvania P, 1978. 275--312.
Denison, David, Ricardo Bermúdez-Otero, Chris McCully, and Emma Moore, eds. Analysing Older English. Cambridge: CUP, 2012.

[ | ѥڡ ]

2012-10-11 Thu

#1263. The LAEME Corpus ɽ (2) [laeme][corpus][representativeness]

ε[2012-10-10-1]˰³The LAEME Corpus ɽꡥϡΤˤƱѥʸˡͿƤ (tagged words) οˤꡤ头Ȥɽͤ롥ޤɽǤ褦

Table 2: Dialectal and Diachronic Distribution of Linguistic Evidence by Number of Tagged Words

 C12bC13aC13bC14aTotal
N0 (0.000%)362 (0.062)0 (0.000)52,883 (9.083)53,245 (9.146)
NEM11,342 (1.948)0 (0.000)3,980 (0.684)2,344 (0.403)17,666 (3.034)
NWM0 (0.000)58,332 (10.019)16,173 (2.778)0 (0.000)74,505 (12.797)
SEM40,082 (6.885)26,722 (4.590)21,921 (3.765)31,408 (5.395)120,133 (20.634)
SWM1,030 (0.177)90,400 (15.527)106,981 (18.375)108 (0.019)198,519 (34.098)
SW1,168 (0.201)2,610 (0.448)46,032 (7.907)30,517 (5.242)80,327 (13.797)
SE0 (0.000)4,043 (0.694)3,199 (0.549)30,561 (5.249)37,803 (6.493)
Total53,622 (9.210)182,469 (31.341)198,286 (34.058)147,821 (25.390)582,198 (100.000)


ľŪǤ褦ˡʬۤ⥶ץåȤɽΤޤǤʰѤˤPDFɤˡ

Dialect/Period Distribution of Tagged Words

ʬۤФϰǤ롥γƥåȤƥȤμʤɤ٤Ĵ٤ȡ˽פ꤬Ƥ롥ĤΥåȤǤϡʬۤΰ찮ΥƥȤˤäƤΤǤ롥㤨СN C14a ȤåȤϡΤΤʤ4ܤ˼Ͽ¿åȤθ95.61% Cursor Mundi Ȥ1ʡΤˤϡɽ魯3ΰۤʤ̸ȿǤ 3 scribal texts [##296, 297, 298]ˤƤ롥ƱͤˡNEM C13b Ǥ #182 Τߤ80.93%θСƤ롥NWM C13b Ǥ #272 Τߤ93.11%SEM C12b Ǥϰۤʤ2ͤμ̻μˤ Trinity Homilies (##1200, 1300) 84.06%ᡤSEM C13a Ǥۤʤ2ͤμ̻μˤ Vices and Virtues (##64, 65) 93.83%롥SW C13b #1600 ϡ69.71%롤
㤬뤳Ȥϡ她åȤɬ⤽θѼɽƤ櫓ǤϤʤषΥƥȤ˸ѼɽƤȤȤ⤷ʤȤȤThe LAEME Corpus λѤκݤˤϡʤؤդɬפǤ롥

[ | ѥڡ ]

2012-10-10 Wed

#1262. The LAEME Corpus ɽ (1) [laeme][corpus][representativeness]

δؿ濴ϽѸηǤ롥λ˴ؿļԤˤȤäƤϡLAEME ԼԤˤСȯ /ˈleɪmiː/ ˤȤ The LAEME Corpus (Text Database) оϡƱ˴ؤ븦ĶġȤơ¤˴ޤ롥LAEME ˤĤƤϡܥ֥Ǥ laeme εǺΤꤢƤȤ櫓ġȤƤβǽõꡤĥ٤#846. HelMapperUK --- hellog ͤαѹϿ޺ CGI ([2011-08-21-1]) #856. LAEME text database Υǡȥƥȵϡ ([2011-08-31-1]) #942. LAEME Index of Sources θġ ([2011-11-25-1]) #1057. LAEME Index of Sources θġ Ver. 2 ([2012-03-19-1]) ɽƤ
繩ˤȤäƻμ줬ʤ褦ˡԤˤȤäƥġθǤ롥Ū The LAEME Corpus ȤäƤ뤦ˡΤȤפȤɤΤ褦ʥѥʤΤΤꤿʤäƤ[2010-11-16-1]ε#568. ѥȱѸ쥳ѥפǼ̤ꡤѥμ礿ħ1Ĥ representativeness ɽˤ롥ϡѥɾΤλɸ1ĤǤ⤢롥˥ѥˤɽγݤ񤷤ˤĤƤϡ#531. OED ΰѥǡ򥳡ѥȤƻȤ뤫 ([2010-10-10-1]) #1243. ٤θ̻ŪΤˡ ([2012-09-21-1]) Ǥ⿨ƤǤ The LAEME Corpus 򶯤Ƥ롥СƤʬۤˤĤƤϡ#856. LAEME text database Υǡȥƥȵϡ ([2011-08-31-1]) ǺΤꤢʬ˲äƻʬޤʤ The LAEME Corpus ΥġʬϤߤ
ޤϡϿƤƥȤοͤ롥ѥ "scribal text" Ȥñ̤ǥƥȤϿƤ뤬Ȼˤäʬ̤ȡФ礬狼롥ʤʬȻʬϤ켫ΤˡʤΤʲǤϡŪʶʬʤȤϤäƤ⤢٤κϤ뤬ˤȤơ7Ĥء4ĤؤʬƤ롥ʤ N (Northern), NEM (North-East Midland), NWM (North-West Midland), SEM (South-East Midland), SWM (South-West Midland), SW (Southwestern), SE (Southeastern) ء C12b 12ȾˡC13a, C13b, C14a ءѸʬˤĤƤϡ#130. Ѹʬ ([2009-09-04-1]) ⻲ȡ

Table 1: Dialectal and Diachronic Distribution of Linguistic Evidence by Number of Texts

 C12bC13aC13bC14aTotal
N0 (0.00%)1 (0.86)0 (0.00)7 (6.03)8 (6.90)
NEM1 (0.86)0 (0.00)5 (4.31)2 (1.72)8 (6.90)
NWM0 (0.00)9 (7.76)5 (4.31)0 (0.00)14 (12.07)
SEM4 (3.45)7 (6.03)14 (12.07)7 (6.03)32 (27.59)
SWM2 (1.72)13 (11.21)17 (14.66)1 (0.86)33 (28.45)
SW3 (2.59)5 (4.31)7 (6.03)2 (1.72)17 (14.66)
SE0 (0.00)2 (1.72)1 (0.86)1 (0.86)4 (3.45)
Total10 (8.62)37 (31.90)49 (42.24)20 (17.24)116 (100.00)


ɽˤоݤȤΤϡThe LAEME Corpus ˼ϿƤ167Ĥ scribal texts ΤȾȤñ̤ǻζʬʤƤ116ĤΤߤǤ롥
ɽͤФ狼褦ˡƥʬۤФ礭Ǥ SEM SWM ؤ۾˸Τ3ʬ2ۤɤ򥫥СƤ뤬 N, NEM, SE ؤǤߤȡC13a C13b 7ۤC12b C14a ؤȤ߹碌Ǥϡ6åȤޤǤ "0" 򼨤˥ѥԻˤ representative γݤ˾ŪȤפƤ롥ʤȤ⡤The LAEME Corpus ѤˤĤƤΥǡ䤽ϡ褯褯դƲᤷʤФʤʤȤȤ
ɽ scribal text οȤ˺Ƥ뤬 scribal text ĹϤޤޤǤ롥ǡƥȿǤϤʤˤʬۤζĴ٤Ƥߤɬפ롥˴Ťɽεϡεǡ

[ | ѥڡ ]

2012-09-21 Fri

#1243. ٤θ̻ŪΤ [frequency][corpus][representativeness]

ε#1242. -ate ưζܹԡ ([2012-09-20-1]) #1239. Frequency Actuation Hypothesis ([2012-09-17-1]) Ǽ夲 Phillips θΤ褦ˡ٤θѲθˤ¿ʴؿ󤻤Ƥ뤬ˡѤʵȤơ٤켫Τ̻ŪѤȤ¤ɤΤ褦˹ͤФ褤ΤȤ꤬롥θǤϤʤäΤ뤤ʬͤˤϡ1--2λ礭ǤϤʤľ롥1λǤϤɤ2ǤϤɤȹͤȡɤޤľΤϤʤϤʤPhillips (225--26) ϡˤĤƼΤ褦˳ڴѤƤ롥

The words' frequencies are based on present-day English, but the general pattern of relative frequencies probably holds for the English in our data base (1755--1993) as well. For example, I would be very surprised if the 3-syllable verbs with CELEX frequencies over 100 --- concentrate, demonstrate, illustrate, contemplate, compensate, designate, and alternate --- were not also much more common in 1755 than those with frequencies of 0 --- altercate, auscultate, condensate, defalcate, eructate, exculpate, expuergate, extirpate, fecundate, etc.


2;λˤƤʤ顤٤100ʾθ0θ٤ȤΤ绨Ĥˤ褦˻פ롥ΤˡPhillips ϼºݤʬϤǤ101ʾ塤10--1001--10ȤӤʬѤƤꡤ绨Ĥپ绨ĤʤޤޤѤ뿵ŤϼƤ롥⤷ä10--100դ٥٥θܺ٤Ĵ٤褦ȤΤǤС2δ֤ˤʤ٤ѲƤǽϤ롥Phillips ʤ餺Ȥ⡤٤Ѥ̻Ū˴ؿï⤬ͤϤ
˻פĤñʲƤϡƻɽǤ礭ʥѥѤɽ뤳ȤǤ롥ƤȤƤñºݤ˿ԤΤϰ֤֤⤫롥ֻٸꤷѸǤСѥѰդɽμưǤѸǤֻ variation 椨 lemmatise Ƥʤ¤ϸФñ̤ǤɽҤޤ夬ŤʤФʤۤɡѥ˴ޤޤƥȤ representativeness Ͽˤʤ롥ӤäݤɽǤ⡤ʤϤۤ褤ƤߤȻפäƤ롥뤤ϡˤäƤϤǤˤ
ʤѤˤ CELEX Ȥñǡ١ϡѸθ֤˴ؤŪʸǤ褯ȤƤΤǤ롥ܺ٤ϡCELEX2 򻲾ȡޤ٤̻֤δطˤĤƤϡ[2012-05-03-1]ε#1102. Zipf's law ȸοաפ򻲾ȡ

Phillips, Betty S. "Word Frequency and Lexical Diffusion in English Stress Shifts." Germanic Linguistics. Ed. Richard Hogg and Linda van Bergen. Amsterdam: John Benjamins, 1998. 223--32.

[ | ѥڡ ]

2011-06-09 Thu

#773. PPCMBE COHA [corpus][coha][ppcmbe][lmode][adjective][comparison][inflection][representativeness]

ܥ֥Ǥⲿ٤夲Ƥ2Ĥ˱Ѹ쥳ѥ PPCMBE ( Penn Parsed Corpus of Modern British English; see [2010-03-03-1]. ) COHA ( Corpus of Historical American English; see [2010-09-19-1]. ) ˤĤơܻ᤬رѸ쥳ѥ٤κǿ˸ΡȤȯɽƤ롥ξԤȤ2010ǯ˸줿ѸΥѥ줾ѼǤ뤳ȡޤԻŪۤʤ뤳Ȥ٤ӤоݤˤŬʤɽϤȤ륳ѥΰŪħ٤뤳Ȥϰ̣
PPCMBE 1700--1914ǯΥꥹѸƥ949,000ǹƤꡤParsed Corpora of Historical English 1ʤƱͤ˹ʸϤ줿Ťб륳ѥȤ³ռǤ롥ͭǥǡꤹɬפ롥COHA 1810--2009ǯΥꥫѸƥ4Ͽ祳ѥǤ롥ϡʸϤϤƤʤCOHA ̵ǥ饤󥢥Ǥ뤿Ȥ䤹󥿡եꤵƤΤǽʥǡǤʤȤ롥
ѥεϤȤط뤬PPCMBE ɽ (representativeness) 񤬤롥PPCMBE ΥѥƥȤ18غ٤ʬषƥǯ10ǯߤǤȤȡȤʤޥܤ¿롥ϡʬ٤ͭյʬϷ̤ФʤȤȤǤꡤѤ˺ݤդפ롥
COHA ΥѥƥȤ Fiction, Popular Magazines, Newspapers, Non-Fiction Books 4绨Ĥ˶ʬƤ롥٤ʬθˤѤǤʤ10ǯߤǤƥޥܤŬڤʥΥƥȤۤƤꡤɽϤ褯ݤƤ롥Fiction ιΨɤλ50%ƤꡤFiction θħä˸áˤѥΤθħ˱ƶͿƤȹͤ졤ʬϤκݤˤϤդפ롥
ܻϡξѥΰʾħ򡤸Ѹˤƻӵ顦ǾˤäƼƤ롥CONCE (Corpus of Nineteenth-Century English) Ѥ Kytö and Romaine ԸˤС19δ֡ӵαФ޷γϡ30ǯߤƬ57.1%67.8%ؤäƤȤƱͤĴ COHA PPCMBE 10ǯߤ˻ܤȤԤǤ1810ǯ64.7%1910ǯ74.3%¤äƤ뤳ȤΤ줿ԤǤ1810ǯ79.4%1910ǯ78.0%ޤɤ줬㤷äȤܡp. 56ˡCONCEƱͤ30ǯߤʬϤľȡPPCMBE ǤͭդѲܴۤѻǤۤɤη̤ǤȤ
ѥϤ줾ȼħäƤ롥褯İѤɬפ뤳ȤǧϢơ[2010-06-04-1]εή˵դäƤӵˡפ򻲾ȡ

2ĤλŦѥ---ɽסرѸ쥳ѥ18桤Ѹ쥳ѥز2011ǯ49--59ǡ
Kytö, M. and S. Romaine. "Adjective Comparison in Nineteenth-Century English." Nineteenth-Century English: Stability and Change. Ed. M. Kytö, M. Rydén, and E. Smitterberg. Cambridge: CUP, 2006. 194--214.

Referrer (Inside): [2017-08-15-1] [2015-09-29-1]

[ | ѥڡ ]

2010-11-16 Tue

#568. ѥȱѸ쥳ѥ [corpus][link][representativeness]

츦ˤ corpus ֥ѥפ͡Ƥ뤬McEnery et al. ʷǤ롥

. . . a corpus is a collection of (1) machine-readable (2) authentic texts (including transcripts of spoken data) which is (3) sampled to be (4) representative of a particular language or language variety.


(1) (2) ˤĤƤϤ褽Դ֤˥󥻥󥵥뤬(3) (4) ˤĤƤϲä "sampled" 뤤 "representative" ȤߤʤˤĤ͡ʰո롥ڤˤƤ뤳ȤǤ
ڤ˱Ѹ쥳ѥˤϡ饤ΤΤǤ롥ʲϡϿɬפʤΤ⤢뤬˥饤ǴؤѤǤѸ쥳ѥ

British National Corpus ʤĤΥ󥿡ե󶡤Ƥ

* BNC ( The British National Corpus )
* BNCweb ̵Ͽ
* BYU-BNC ̵Ͽ

BYU Corpora Brigham Young University, Mark Davies 󶡤Τ¾Υ饤󥳡ѥ

* COCA ( Corpus of Contemporary American English ) ̵Ͽ
* COHA ( Corpus of Historical American English ) ̵Ͽ
* TIME Magazine Corpus of American English ̵Ͽ

Cobuild Concordance and Collocations Sampler

¾ܥ֥Ǥϥѥطε򤤤ȷǺܤƤΤǡͤˤ줿

hellog Υѥν󵭻: [2010-09-15-1]
hellog ΥѥϢ: corpus
hellog BNC Ϣ: bnc

McEnery, Tony, Richard Xiao, and Yukio Tono. Corpus-Based Language Studies: An Advanced Resource Book. London: Routledge, 2006.

[ | ѥڡ ]

2010-10-10 Sun

#531. OED ΰѥǡ򥳡ѥȤƻȤ뤫 [oed][corpus][representativeness]

OED (2nd ed. CD-ROM) ˱Ѹ쥳ѥȤѤȤȯۤäŻǤǤƤ鹭ͭƤºݤ¿θ OED ѥȤƳѤƤ롥⤽⤬ѥȤԤޤ줿櫓ǤϤʤ OED ν򥳡ѥȤߤʤƸ椹뤳Ȥϡɤ줯餤ʤΤƻˤĤΤ뤳Ȥϸ漫ȤƱ餤פȻפΤǡΥơޤ˴Ϣ Hoffmann ʸޤȤƤߤʻ伫ȤƻȤƤ OED ħ褯򤻤˸˻ȤäƤ餤ΤǡʬΤ˺ϿȤĤǤսν񤫤줿ʸ򻲹ͤˤƤޤ
Hoffmann OED ν򥳡ѥȤѤ뤳ȤǤ뤫ȤФơ4Ĥδ饢ץƤ롥ƴȡб Hoffmann η󤹤롥

(1) Selection criteria for the quotations
"a collection of pieces of language that are selected and ordered according to explicit linguistic criteria in order to be used as a sample of the language" (19; cited from Sinclair) Ȥ̩ʥѥ˾Ȥ餻СOED ν򥳡ѥȸʤȤϤǤʤΤˡġθФ첼ǼƤ㷲θФܤŬڤʥѥˤʤʤȤȤϸθü٤η֤̣åפ뷹뤫Ǥ롥äˤ븫ФܤΤǤʤСΤȤ OED ϳƻαѸɽƤȹͤ졤ѥȤƳѤ뤳ȤǤ롥

(2) Representativeness and balance of the quotations
OED ϼºݤ˲餫ŵ򤫤Ƥ "true quotations" (20) Ǥ롥ԼԤˤäƺ줿ʤǤϤʤϤƾʤޤŵΥ¿ˤ錄ꡤüʸغʤ˸¤ʤɤиʤΤǡ˴ؤƤ "representative" ȸäƤ褤ƥ뤬츦ˤȤäŬڤʳʬۤƤ櫓ǤϤʤΤǡ"balanced" Ȥϸʤ㤨 Shakespeare 1ͤ33,000󶡤Ƥʤɤ󤲤롥OED 򥳡ѥȤƸΩƤˤϡ"balance" դפ롥

(3) Reliability of the data format
ʸΰάƤ褦㤬ʿѤ20?25%ۤɤ롥ۤȤɤξάǤʸι¤ƤʤˤŬڤʾάʸι¤ѲƤޤäƤʸ⤢롥ʾι¤Ĵ٤뤿 OED ѤˤϡդɬפǤ롥

(4) Quantification of the results
1ǯ򥰥դ˥ץåȤȡ174000ۤ뾮ԡ1910000ۤԡǧ뤬20ˤϷ㸺롥ǡοϻˤ餺13٤Ȱǡ20㤬ĹʤΤܤαޤ٤Ǥ롥240ۤʽǤ180ۤɤäˤȤȾ嵭ʿѸ׻ơOED ˴ޤޤ3300?3500ȿꤵ롥OED 򥳡ѥȤѤˤϡ19ä¿ȤʤɤդƸ̤᤹٤

Ǹ Hoffmann ηѤ (26) OED νϸѲη绨ĤŪɽ魯ѥȤƸѲˤȤäͭѤǤ롤ȤQŪʷŪʿФƤƻͤˤʤä

Although the OED quotations database is not a completely balanced and representative corpus, it can nevertheless provide the linguist with a wealth of useful information. The data it contains chiefly represents naturally occurring language, and the time-span covered is unmatched by any other source of computerized data. Even though over 20 per cent of all its quotations have been shortened, the large majority of these deletions is unlikely to distort the results of many diachronic studies of linguistic features. Given the nature of the data, normalized frequency counts might suggest an inappropriate level of precision, but tendencies in the development over time can nevertheless be expressed in quantitative terms. (26)


The Oxford English Dictionary. 2nd ed. CD-ROM. Version 3.1. Oxford: OUP, 2004.
Hoffmann, Sebastian. "Using the OED quotations database as a Corpus --- A Linguistic Appraisal." ICAME Journal 28 (April 2004): 17--30. Available online at http://icame.uib.no/ij28/index.html .
Tanabe, Harumi. "The Rivalry of give up and its Synonymous Verbs in Modern English." Language Change and Variation from Old English and Late Modern English: A Festschrift for Minoji Akimoto. Ed. Merja Kytö, John Scahill, and Harumi Tanabe. Bern: Peter Lang, 2010. 253--75.

[ | ѥڡ ]

Powered by WinChalow1.0rc4 based on chalow