ɸˤĤƤϡʲεޤ͡ʵ˼夲Ƥ
#307. ѥѤ ([2010-02-28-1])
#367. ѥѤ (2) ([2010-04-29-1])
#428. The Brown family of corpora Ѿա ([2010-06-29-1])
#1280. ѥɽ ([2012-10-28-1])
#2584. ˱Ѹ쥳ѥɽ ([2016-05-24-1])
#2779. ѥϱѸ˸˻Ȥ뤱ɤ ([2016-12-05-1])
ѥѤѸʻˡ˸ϤޤޤˤʤäƤƤꡤسǤ뤵褦ˤʤä餳ѤˤäǧƤȤǤ롥ݤϤ褽֤Ȥʤ뤬ϱѸγ路 Fischer et al. (14) ꡤ4Ŧ褦
(i) there can be tension between what is easily retrieved through corpus searches and what is thought to be linguistically most significant; a historical syntactic case in point involves patterns of co-reference of noun phrases . . . ; these have been largely neglected because they involve information status, which is currently not part of any standard annotation scheme;
(ii) when a data search yields large numbers of hits, there may be a temptation to interpret corpus results merely as numbers, which is a severely reductive approach; in cases of grammaticalization, for example, changes in frequency may act as tell-tale signs . . . , but an exclusive quantitative focus will mean that one is ignoring the changes in meaning and context that form the core of the process;
(iii) the substantial amounts of data that can be collected from a corpus can also blind researchers to the dangers of making generalizations about the language as a whole on the basis of a partial view of it; this is a particularly relevant problem for diachronic research, because we only have very incomplete evidence for the state of the language in any historical period . . . ;
(iv) trying to achieve greater representativness by collecting and comparing data from various corpora can also be tricky: principles guiding text inclusion vary widely, there is little standardization in user interfaces, and they can require a significant time investment to learn to operate.
4θդĶСΤ褦ˤʤ롥
(i) ѥǿԤ䤹꤬Ūˤɬ̣ΤǤϤʤ⤷ʤդ٤
(ii) ŪʴŻ뤹븦ˤΩŪʴᤴƤޤ
(iii) ʥѥǤäȤƤ⡤ representative Ǥ櫓ǤϤʤʤ˸ؤˤ "bad-data problem"
(iv) ѥԻԤ䥤եԤΰտޤĤǡˡƽϤ٤
Fischer, Olga, Hendrik De Smet, and Wim van der Wurff. A Brief History of English Syntax. Cambridge: CUP, 2017.
#1264. ˸ؤθ³ȡιؤƻ ([2012-10-12-1])#2865. Ĥ䤹ڵä䤹ڵؤΥҥȡ ([2017-03-01-1]) Ǽ夲Ƥ˸ؤˤϻθ³ȤȤ⤷꤬롥̤Ȥ˾ۤɤΤΤĤäƤƤʤΤ¤Ǥ롥 (29) ϡ˸ء٤ΤʤΡָűѸαŪŪפȤˤƼΤ褦˽Ҥ٤Ƥ롥
˸ؤǤϸ¸ǽפǤ뤳ȤϤޤǤʤθоݤȤΤǤСʸϿ˲äơüԤθľѡʤʤɤθ쿴ŪޤƼ¤˭٤ʻȤΤǤ뤬˸ؤǤϤϴñˤʤʤŤλȤȤ¿ʬʤΤǡμθ³Τ褦˻פ뤬ºݤϡʾ˸ΤǧʤФʤʤä˰ŤиŤۤɸΤꡤѸˤǤϡäOEθ³ˤĤƤϤ褯ǧǡʤƤʤƤϤʤʤ
Ūˤɤ줯餤Τָ뤿ˡűѸѸλˤĤƻϤƤս⤷Ƥ
űѸμܤ˴ޤޤ300ǡʸ2000Ǥ롥ʬŪˤϥˤ˲Ǥ롥̤ϥΥޥʹߤ200ǯ֤˽줿Ѹλ⾯ʤ(29)
Ȥ850ǯλǻĤäƤΤϡ4ĤΥƥȤ35ۤɤˡΧʸļʤɤûŪʸȾǤ롥(29)
űѸ9䤬ȥǽ줿Ǥ롥(31)
ɮ (authorial holograph) ѸǤ Ayenbite of Inwit 1340ǯˡ Hoccleve (1370?--1450?) νʪ15Ρեε²μˤʽ Paston Letters 䤽¾ƱνʽۤɤǤ롥(34)
Ѹˤˤ礭ʸ (philology) ʸɾ (textual criticism) Υץʬǽ뤵ʤǤ롥
Ϣơ#1264. ˸ؤθ³ȡιؤƻ ([2012-10-12-1])#2865. Ĥ䤹ڵä䤹ڵؤΥҥȡ ([2017-03-01-1])#1051. Ѹ˸оݤȤʤ (1) ([2012-03-13-1])#1052. Ѹ˸оݤȤʤ (2) ([2012-03-14-1]) ⻲ȡ
2ϡѸ˳ѡ ԡˡ˸ءīоȸإȯŸԡ3īҽŹ2018ǯ22--46ǡ
ˤ椹ۤͣˡϡ¸˰͵뤳ȤǤ롥Ȥ¸ϼŪˤŪˤФ꤬ꡤѥؤѸǤ "representative" Ǥ "balanced" Ǥʤ
ޤʪŪʾ郎롥ߤ뤿ˤϡĹ֤Ѥ̺ƻǵƤ뤳ȤɬפǤʡ#2457. ̺Ƚƻ (2) ([2016-01-18-1]) ȡˡˡǥδ顤äաŪʸոˤϡĤơߤޤǽ㤤Ȥ餫δϡʸˤϤäѤɤ߽ǽϤΤ륨ؤθո乥ߤ꤬ȿǤ졤ʳؤθưϿ˻Ĥ뤳ȤϤۤȤɤʤˡƤȤ顤㤨ȿŪʽʪϡˤǽ⤤
¸ʸϡ͡ʱ̿ȴĤäƤȤ̣ˤΡֶפθƤΤϳΤĤ䤹ΤˤϤĤξ郎Ȥ嵭εǰƬ֤ȡΡɬפθƤȤ롥˸ؤʸؤˤơɤΤ褦ʡ־פɬפΤƤȤϽפ
ΤδŪʥҥȤȤơɤΤ褦ʲФפߤޤĤ䤹ͻ벽 (taphonomy) Ȥʬθͤˤ㤨СಽФλĤ䤹ĤˤϡˤäƷޤȹͤåɤϡֿಽеϿζкߡפꤹ (73--75) ǡΤ褦˽Ҥ٤Ƥ롥
ǯ֤⡤ؼԤϡ700?600ǯʹߤ¸οΤͳ褹벽ФƤο¿褦˻פ뤫⤷ʤʬϸߤ˶ᤤǯΤΤǯŪкߤΤۤˤ⡤ಽФˤϤޤޤФ꤬롥Τ褦Ф꤬뵡餫ˤȤϡֲءפȤФ롥䲼ܹ뤤ϻͻ礭ʹϤ褯Ĥ뤬ǹϾסؤιʤɤ̩ưפ»ΤǻĤˤĤޤꡤλĤ䤹礭ȴ椵㤹롥ǹΤ褦ʷڤϱˤή졤Ф˱Ф졤ʤˤιȰ˲Фˤʤ롥ŤƬܹϹή졤δδ֤˰ääơۤΦ緿ưʪιȰ˲Фˤʤ롥
ФλĤ⤦ĤװϡῩưʪΤΤɤʬफǤ롥ҥ祦ϥμΤʤΤǡ⤷ΤǤῩưʪƱäΤʤ顤ಽФμȯ뤳ȤʤϤǤ롥ºݡβФϤޤ껺ФƤ餺ΤοʲˤĤƤϤ褯狼äƤ뤬οʲϤ褯狼äƤʤΤ礭⡤ФĤ뤫ɤ˱ƶ롤Τ礭ʼϲФȤƻĤ䤹ƱǤΤ礭ʸΤϻĤ䤹Τ褦ФϿಽФ⳺롥
ĶǤϡۤ١Фˤʤ䤹ȯ䤹äơǯ䤢ϰͳ褹벽Ф¿Ȥäơǯ夢뤤ϰ¿θΤǤȤϸ¤ʤǯϰξۤ٤ƲвŬƤǽΤƱͤˡǯϰ褫ಽФĤʤƤ⡤˿बǤʤäȤˤϤʤʤ־ڵΤʤΤϡ¸ߤʤڵǤϤʤפȤʸ⤢롥äơΤμϡǸŤβФȯǯǿβФȯǯޤ¸ƤȤˤʤ롥ĤޤꡤмεǤǯϼºݤ٤ƾ˹ˤʤ롥
Ʊͤ¤ϲȯפŪʬۤˤƤϤޤ롥ϡФȯפ깭ϰϤ¸ƤϤޤδĶϸߤȤϰäƤϤߤǤϸĶˤϽߤ䤹ä⤷ʤդ⤢롥ˡФȤ¸ƤĶϷ褷¿ϤʤھǤϡä˻ĤʤȤ˿ӴĶǤϡ٤⤯ھʤΤǡФϻĤʤȹͤƤǶǤϡɬ⤽ǤϤʤȤ餫ˤʤäȤϤųؼԤдȿȯȴäƤ⡤ƤϤƤޤд路ȯǤʤΤ롥
ϡνФѸȤʤäʸؾȤϢ˸ˤˤĤƤϡ#1051. Ѹ˸оݤȤʤ (1) ([2012-03-13-1])#1052. Ѹ˸оݤȤʤ (2) ([2012-03-14-1]) Ȥ줿ϢơʸȯˤĤƤιͻ⻲ͤˤʤʡ#1834. ʸǯɽ ([2014-05-05-1])#2389. ʸηϤεȯã (1)([2015-11-11-1])#2457. ̺Ƚƻ (2) ([2016-01-18-1]) ȡˡ
СʡɡåɡˡϾ ͪˡˡؿοʲȻǤˤõ١ǡ2014ǯ
˸ؤȥѥѤȤ褤ȸ롥㤨бѸˤǹͤСŪʥƥȤͭ¤Ǥ뤫顤٤ƤŻŪ˳ǼСʸ̤ŪĴ뤳ȤǤ롥αѸ˵Ҹ⡤ŻŪǤʤäΥƥȷĴоݤȤƤǤϡƱ֥ѥäȤ⤢롥ߤǤϡޤޤ¿θԤŻҥѥȤδϢġѤƸ椷Ƥ뤷Ĭήϴְ㤤ʤȯŸƤ
դ٤Ȥ⤢롥Ѹ쥳ѥؤ路ɥ (197) Τ褦˽Ҥ٤Ƥ롥
1990ǯʹߡ˸ؼԤˤȤäƥѥ˽פʥġȤʤޤȤʸʤФʤʤä⡤ʤΨ褯ܤõˤʤޤȤϤ˸ؤǤϡΥѥȤäʾˡʸΩ֤ꡤɬפޤʤŤǤϡΡָʸפɬ⸶ܤǤϤʤʣ̤Ǥ뤫⤷줺ˤϤʣ̤⡤ʣμ̻Ȥ˼DzԤƻĤΤΰĤˤʤǽޤˡ顤μŪʥѥؤϡɴǯθθȤۤˤޤŤαѸ쥳ѥʤСѸˤμȤȤδϢǹԤΤ褤Ǥ礦
Ѥ˽Ҥ٤Ƥ̤ꡤȤ櫓ŤθоݤȤˤϡդפ롥θΤ褦ΤޤޤǤȡˤ֤Ĥ뤫㤨СѸˤǤСѸǤֻꤷƤʤñ褦ȤƤ⡤ɸŪֻʤΤǡͤ͡ʰ֤ǸƤߤɬפ롥ޤʸˡŪʥդʤƤˤ⡤ѸȸűѸ졦ѸȤǤʸˡƤۤʤäƤ뤷μ㤷ƤΤǡˤäƻ˸ŤʸˡμɬܤȤʤ롥פˡ˸Ť˽ϤƤ뤳ȤΤǤ롥
ˡѤǸڤƤ̤ꡤָʸפοԲķǤˤ⤫餺ѥɽŪѤ뤳Ȥ˴ƤȡƤοߤ˰ռʤʤ꤬Сʸɾ (textual criticism) ʸŪʥץ±ˤʤäƤޤ롥ѥΤȤ櫓ǤϤʤ椹ԤΤȤռŪ˼ФƤФ褤ȤȤǤϤΤñǤϤʤ
⤦1ġŻҥѥʤ餺Ȥˤʤ뤬ѥɽ (representativeness) ꤬ˤ롥ŤξˤϡߤƥȤ϶ĤäƤȤפΤʤ𤬤ꡤɽϤȤ˿
ŤθѥĴȤȤϡŪˤŪˤ⡤츫ۤɤ䤹ʤȤȤƤϢơʲε⻲ȡ
#568. ѥȱѸ쥳ѥ ([2010-11-16-1])
#363. Ѹ쥳ѥȯŸ3 ([2010-04-25-1])
#368. ѥϸβǽ ([2010-04-30-1])
#1165. ѹǥѥ椬ˤʤäطʡ ([2012-07-05-1])
#307. ѥѤ ([2010-02-28-1])
#367. ѥѤ (2) ([2010-04-29-1])
#1280. ѥɽ ([2012-10-28-1])
#2584. ˱Ѹ쥳ѥɽ ([2016-05-24-1])
ϡɥȡˡ 翹 ʸҡ ޤߡ ɹˡرѸ쥳ѥѤ츦١罤ۡ2016ǯ
#1917. numb ([2014-07-27-1]) εǡѸˤ nimen ȸťΥɸѸ taken ζˤĤĴ Rynell θ˿줿̤˸ťΥɸѸ줬Ѹ椤ˤƱѸ˿ƩƤäݤˤϡδϰδθ롥ΤȤʤ顤οƩˤϤ٤λ֤ΤǡΤۤƩٹ礤ϸȤʤޤťΥɸαƶ the Danelaw ȸƤФ륤ˤƺǤǤꡤؤϡξ⤬֤ޤʤŤƤäȹͤΤǤ롥
Τ褦˸ťΥɸθŪƶζˤĤƤϡϰδ֤̩ܤߴطꡤʬۤΤǤȤ롥ºݤˡ#818. ɤ˻ĤťΥɸ̾ ([2011-07-24-1]) #1937. Ϣ -son ˤΤϸťΥɸͳ ([2014-08-16-1]) ˼ʬۿޤϡΤʬۤűѸȸťΥɸѸ줬礹륱Ǥϡ̤˾嵭ʬۤǧ뤳Ȥ¿褦Rynell (359) ۩"The Scn words so far dealt with have this in common that they prevail in the East Midlands, the North, and the North West Midlands, or in one or two of these districts, while their native synonyms hold the field in the South West Midlands and the South."
ϰ츫ۤñǤϤʤȤˤαդɬפ롥Rynell (359--60) Ͼʸ³ơΤ褦âդäƤ롥
This is obviously not tantamount to saying that the native words are wanting in the former parts of the country and, inversely, that the Scn words are all absent from the latter. Instead, the native words are by no means infrequent in the East Midlands, the North, and the North West Midlands, or at least in parts of these districts, and not a few Scn loan-words turn up in the South West Midlands and the South, particularly near the East Midland border in Essex, once the southernmost country of the Danelaw. Moreover, some Scn words seem to have been more generally accepted down there at a surprisingly early stage, in some cases even at the expense of their native equivalents.
äդ٤ϡ¸ѸƥȤʬۤФäƤǤ롥СѸ쥳ѥϰ˴ؤɽ (representativeness) 礤ƤȤRynell (358) ˤС
A survey of the entire material above collected, which suffers from the weakness that the texts from the North and the North (and Central) West Midlands are all comparatively late and those from the South West Midlands nearly all early, while the East Midland and Southern texts, particularly the former, represent various periods, shows that in a number of cases the Scn words do prevail in the East Midlands, the North, and the North (and sometimes Central) West Midlands and the South, exclusive of Chaucer's London . . . .
ťΥɸθŪƶϡѸᤤǡ٤ˤǴѻ롤ȤȤϳȤƽҤ٤뤳ȤϤǤΤΡ줬Ѹ쥳ѥλʬۤȸ˰פƤ¤ƨƤϤʤʤĤޤꡤ嵭γŪʬۤϡޤ¸ƥȤλ֡ŪʬۤʿԤƤ뤿ˡȤˤ˶ĴƤ뤫⤷ʤΤ䤹Τޤޤ䤹ʤꡤˤΤ줿ޤޤˤ빽¤Ū꤬ˤ롥
ϡťΥɸθŪƶˤȤɤޤ餺Ѹ̡ŤѲ̤ѻݤˤͿǤ (see #941. ѸθѲϤʤ̤ŤΤ ([2011-11-24-1])#1843. conservative radicalism ([2014-05-14-1]))
ϢơѸ쥳ѥ A Linguistic Atlas of Early Middle English (LAEME) ɽˤĤơ#1262. The LAEME Corpus ɽ (1) ([2012-10-10-1])#1263. The LAEME Corpus ɽ (2) ([2012-10-11-1]) ⻲ȡ
Rynell, Alarik. The Rivalry of Scandinavian and Native Synonyms in Middle English Especially taken and nimen. Lund: Håkan Ohlssons, 1948.
ѥɽ (representativeness) ѹ (balance) ˤĤƤϡ#1280. ѥɽ ([2012-10-28-1]) ¾εǰäƤŤαѸΥѥˤϡѸ쥳ѥ˴ؤ꤬ΤޤƤϤޤΤΤȤʤ顤˾褻Ƥ˺꤬¿ΩϤ롥
ޤ˱Ѹ쥳ѥȤȤơѥθʤ콸ΥƥȽΤΤΤˤζˤ긽¸ƤΤ˸¤Ȥ롥ʸܡܡʤɤ˵ƸߤޤĤꡤ¸ƤΤ٤ƤǤ롥ޤλ¸ߤ뤳ȤϤ狼äƤƤ⡤Ū˥Ǥ뤫ɤǤ롥Ūˤϡ뤤Żҷ֤ǽǤƤ뤫ɤˤäƤλƥȤ콸Ȥʤꡤ褯ԻԤˤäΰѥʼȤŻҷ֡ˤؤԻ뤳Ȥˤʤ롥Ω˱Ѹ쥳ѥϡäơ褦䤯˽ФΤǤꡤλŪɽãƤ븫ߤϡǰʤ
ޤ˱ѸȤҤȤ˸äƤ⡤ºݤˤϸѸƱͤ͡ lects registers ʬ졤ζʬ˱ƥѥԻ륱¿Τˡ̣ѥѥȸƤǤ褤 Helsinki Corpus Τ褦̻ѥ䡤òƤ뤬ۤʤޤ Penn Parsed Corpora of Historical English ⤢뤷ȤˤäƤ̻ѥȤƤѤǤ OED ΰʸʤɤ롥̾ϡԻŪ֤˱ơ꾮ϰϤΥƥȤ˹ʤäԻ륳ѥ¿űѸ쥳ѥѸ쥳ѥʤɻˤäƶڤ뤳 (chronolects) ⤢ ꥹѸ䥢ꥫѸʤɤ (dialects) ξ⤢뤷Chaucer Shakespeare ʤκ (idiolects) ξ⤢Ҳ (sociolects) ̤Ȥ⤢뤷Ѱ (registers) ˱ƥѥԻȤȤ⤢롥ѰȤäƤ⡤äξʥˡΡäդդˡʷˤʤɤ˱ơ̶ʬ뤳ȤǤ롥ʬޤ٤ƤޤȡҤΤ褦˸¸ƥȤ̤ͭ¤ǤꡤƤ˾ʤäʬۤФäƤ櫓顤ɽѹդݤĤȤʤΤȺȤʤ롥
(chronolects) μ濴ˤƶǯŪ絬Ϥ˱Ѹ쥳ѥԻѤƤߤȡƻαѸμʸˡϿޤΤ褦ʻͻԻȴϢŤԻ줿ΤĤ뤳Ȥ狼롥ϡƻδѥȤư֤ŤȤäƤ褤⤷ʤ㤨СDictionary of Old English Corpus (DOEC), A Linguistic Atlas of Early Middle English (LAEME), Middle English Grammar project (MEG) Ǥ롥ˤĤƤϥѥȤϥƥȡǡ١Ȥ٤ EEBO (Early English Books Online) ѲǽȤʤäƤƤ뤷ꥫѸˤĤƤ Corpus of Historical American English (COHA) λߤ⤢롥
˵ɽŪ̾ʤΤΤۤ͡ڤ˱Ѹ쥳ѥԻδ褬³ȸƤ롥ˤĤƤϡ#506. CoRD --- Ѹ˥ѥξ ([2010-09-15-1]) ǾҲ𤷤Helsinki ؤ VARIENG ( Research Unit for Variation, Contacts and Change in English ) ץȤ CoRD ( Corpus Resource Database ) Ȥ줿
ϡŪˤСҤΤ褦ˡ˱Ѹ쥳ѥɽŪ˲褹ΤϺǤϤ뤬Ǹ̥ѥԻϳ趷褷ƤꡤΤʤǿⴶϤ롥ҤȤĤϥեŪ¦̡Ѽ¦ΥѥФ٤˴ؤ⤢褦˻פ롥Żҥѥλ夬褹ˤ⡤˱ѸθԤϡ֤פȤʤΤˤ֥ѥפԻѡʬϤ˹ԤʤäƤΤǤ롥⡤䤿ռƤۤɤǤϤʤäΤΡ٤ϥѥɽѹդȤäͤƤΤǤꡤ˻ʤȤƤ⡤ͭפʸ³θѤƤŻҥѥȤʤäɽѹդ꤬ΩäƼ夲褦ˤʤäΤˤ꼫Τڤǡͤ³ɬפϤΤΡ餫ˤ츽ݤΤΤ˾ơѥ˻Ȥʤʤ餽θ³Ƥ椯ȤפʤΤǤϤʤ
ε#2520. Ѹ134 "such" ΰֻ ([2016-03-21-1]) ³ϽѸ쥳ѥ LAEME "such" ΰֻФƤߤ (see #1262. The LAEME Corpus ɽ (1) ([2012-10-10-1])) θϡѸǤϷƻ졤졤³ȤѤ졤ƻξˤ϶ޤ⤹ΤǡΤȤ͡ʷ֤롥ե٥åȽ˰褦ʤäοͤʸڤ١ˡ
hsƿucche (1), schilke (1), schuc (3), scli (1), scuche (1), sec (1), secc (1), secche (1), sech (2), seche (1), selk (1), selke (1), shuc (1), shuch (1), siche (1), silc (1), silk (3), sli (1), slic (5), sliik (1), slik (3), slike (1), slk (1), sly (1), soch (5), soche (1), solchere (1), suc (2), sucche (2), such (51), suche (1), suecche (1), suech (1), sueche (1), sueh (1), sug (1), suic (1), suicchne (1), suich (12), suiche (3), suilc (14), suilce (1), suilch (1), suilk (1), suilke (2), sulch (1), sulche (1), sulk (1), sulke (1), suuche (1), suweche (1), suwilk (1), suyc (1), suych (4), suyche (1), svich (2), sƿche (1), sƿic (3), sƿicche (1), sƿich (1), sƿiche (14), sƿichne (1), sƿilc (30), sƿilch (22), sƿilche (1), sƿilcne (1), sƿilk (14), sƿillc (10), sƿillke (2), sƿi~lch (1), sƿlche (1), sƿuc (4), sƿucch (1), sƿucche (4), sƿucches (1), sƿuch (1), sƿuche (1), sƿuchne (1), sƿuilc (1), sƿulc (8), sƿulce (1), sƿulche (9), swch (1), swecche (1), swech (1), sweche (2), swich (5), swiche (1), swics (1), swil (1), swilc (5), swilce (1), swilk (2), swilke (2), swilkee (2), swlc (1), swlch (1), swlche (1), swlchere (1), swlcne (1), swuche (2), swuh (1), swulcere (1), swulch (3), swulchen (1), swulchere (1), swulke (1), swulne (1), zuich (10), zuiche (14), zuichen (3), zuych (10), zuyche (2)
ʸȾʸζ̤ϤĤˡ113ֻʸڤ롥Τʤ٤ˤƥȥå5ֻȴФȡsuch, sƿilc, sƿilch, sƿilk, sƿiche Ȥʤꡤ5369ĤΤ131 (35.5%) 롥
θѸ줫134ȹ碌ʣֻȡѸΤȤ247ΰֻ뤳Ȥˤʤ롥ѤϿޤ䥳ѥɬŪǤϤʤΤǡϹʿͤȻפ롥㤨СMED swich (adj.) ˷ǤƤֻäСϤ⤦
#1730. AmE-BrE 2006 Frequency Comparer ([2014-01-21-1]) ǡ2006ǯνեƥȤԻƳѼ拾ѥҲ𤷡˴ŤӥġΥġʤ鵤ŤΤƱˡԻ졤ϤƱ100٤ the Brown family of corpora ʡ#428. The Brown family of corpora Ѿա ([2010-06-29-1])ˤϢȤСľ50ǯ֤ۤɤ̻ŪʱƴӤưפ˲ǽȤʤ롥
ǡεǾҲ𤷤 Professor Paul Baker - Linguistics and English Language at Lancaster University ˤ AmE06 BrE06 ˲äơեꥫѸɽ Brown (1961), Frown (1992)եꥹѸɽ LOB (1961), FLOB (1991) ɽФ碌ƥǡ١ѤλϡAmE-BrE 2006 Frequency Comparer ȤۤƱʤΤǡμ ([2014-01-21-1]) Ȥ줿ϤɽǤϡθиƥȤοٽ̤ϾʤƤꡤ100٤ɽˤȤɤƤΤǡAmE06 BE06 ˤĤԤξɬפʾˤϡAmE-BrE 2006 Frequency Comparer ɤ
#1415. shew show (1) ([2013-03-12-1]) ȡ#1416. shew show (2) ([2013-03-13-1]) ǰä#1637. CLMET3.0 between betwixt ʬۤĴ ([2013-10-20-1]) ǾҲ𤷤 The Corpus of Late Modern English Texts, version 3.0 (CLMET3.0) ˤˬŪˤϡƱդѥ "\bshow(s|n|ed|ing)?_VB" "\bshew(s|n|ed|ing)?_VB" Ǹơ3ʬȤٿ٤ʲη̤Ф
| shew | show | ||
|---|---|---|---|
| 1710--1780 | 335 | 1,545 | 10,480,431 |
| 1780--1850 | 159 | 3,100 | 11,285,587 |
| 1850--1920 | 92 | 5,118 | 12,620,207 |
Why, you have shewn your wit upon the subject, and I mean to show your courage;
Mr. Wright, as well as Nadin, professed they were perfectly satisfied of this, and appeared to shew to me all the polite attention that they were capable of showing.
Assuredly I did not show him the face which I shewed Folderico.
ѥˤȤäɽ (representativeness) ̿Ǥ뤳Ȥϡѥ ([2010-11-16-1]) 餫Ǥ뤷ε#1279. BNC ζߤȼߡ ([2012-10-28-1]) ǾҲ𤷤 Leech Ȥ櫓ĥƤǤ롥McEnery et al. (13) ϡɽˤĤơLeech ͤˤʤ "a corpus is thought to be representative of the language variety it is supposed to represent if the findings based on its contents can be generalized to the said language variety" ȽҤ٤Ƥ롥
ɽŪ˹ͤƤߤ褦㤨 BNC åȤȤ褦ʡ奤ꥹѸȤŪѼϿ륳ѥ (general corpus) ɽϤɤΤ褦ˤΤäդȽդγͤȡ줾50%Ĥ˳꿶뤳Ȥϡ奤ꥹѸɽ«ƤLeech ɽǤ "impressionistic" Ȥʤ餶ʤνִ֤˹ԤʤƤ븽奤ꥹѸΰŪʬäդˤƤǤϤʤ⤷ȤСäեѥγ㤨80%ۤɤꤹۤɽݤǤΤǤϤʤΤȤʤ븽奤ꥹѸľܤĤळȤǤʤʾ塤ɽεϹԤͤޤäƤޤ
ѥä˰̥ѥˤɽȤˡ balance sampling Ȥ2Ĥγǰʬƹͤ뤳Ȥ롥McEnery et al. (13) Ǥϡ"the representativeness of most corpora is to a great extent determined by two factors: the range of genres included in a corpus (i.e. balance . . .) and how the text chunks for each genre are selected (i.e. sampling . . .)" Ƥ롥
balance ȤϡBNC ѸǤȤ domain genre Ȥʬ˴ؤΤǤ롥㤨С奤ꥹѸΥѥɸ֤ʤ⡤ꥹοʹαѸѥϡrepresentativeness 롥奤ꥹѸˤϽդǤʤäդ⤢뤷ԤˤĤƤϿʹѸǤʤʸرѸ⤢ŻҥѸ⤢뤷㤤ʪѸ⤢СѸ⤢롥Τ domain genre θ줿Ȼפ̤ƤĤ text domain ΤʹѸ˸¤äƤ⡤֥ɤ⤢й⤢롥1ĤοʹǤ⡤Ҳ̡ݡ̡ʤɤ̤ɬפϤʤΤҲ̤Ǥйȹݵζ̤ϤɤŪˤϤɤޤǤʬ롥äդǤƱͤ˺ʬ䤷ʤƤСĿ (idiolect) ˸Ŀˤ register ̤θ졤ʤɤΥȥؤȽ夷ƤޤºݤΥѥϡQŪʥ٥Ŷ뤳Ȥˤʤ뤬־QŪפ "impressionistic" ϤۤƱ
sampling Ȥɽ뤿μˡǤ롥ΤθŪħƸ褦ˡ̤ˤƹθäʤ顤ѥ˳ domain ۤ뤿ȼǤ롥ˤϡsampling unit ȤƲꤹ뤫ŵŪˤϡܡʹʤɤʤȤƤñ̡ˡΤ褦ñ̤ꥹȲȤϰ (sampling frame) ɤޤǤꤹ뤫ǯؤθ䡤٥ȥ顼ܤؤθʤɡˡɸܼϴʥˤ뤫٤ηϲäǤΥˤ뤫ɤۤ뤫ʤɤΡŪŪ꤬ޤޤ롥
ɽ˴ؤ⤦1ĤγǰȤơclosure 뤤 saturation ȸƤФΤ⤢롥McEnery et al. (16) ˤС"Closure/saturation for a particular linguistic feature (e.g. size of lexicon) of a variety of language (e.g. computer manuals) means that the feature appears to be finite or is subject to very limited variation beyond a certain point." Ƥ롥ʿСʾ女ѥεϤ礭Ƥ⡤ùγѤʤȤϤãСΥѥ saturated Ǥȹͤ롥ɽλɸȤƤϡbalance saturation ΤۤƤȤŦ⤢뤬saturation ϼȤƸäǰƬˤꡤ¾θܤؤαѤϻߤƤʤΤǤ롥
ɽϡ女ѥ̿ǤȤϤäƤ⡤ԤȤ餤Ϥ롥ݤ뤿ʤˡʤ٤ƤΥѥԻԤΩϤƬˤѥϼԻƤ롥Ṳ̄ˤơҤԻȻѤ³Ƥ椭Υϥ٤ʳˤΤ⤷ʤ
McEnery, Tony, Richard Xiao, and Yukio Tono. Corpus-Based Language Studies: An Advanced Resource Book. London: Routledge, 2006.
108--114֤ˤ錄ꡤΩرѸ춵鸦ˤŤǡLancaster ̾ Geoffrey Leech θֱ줿ϡ2ܤ "The British National Corpus: Both a Triumph and a Failure" ꤹֱΤߤλääİ˹ԤäBNC ԼԤκäʤɡ⤷ää
̾ˤ "triumph" "failure" ˤĤơLeech Ϥ줾켡Τ褦ʹܤƤ
A triumph:
It has been claimed that the BNC is the most widely used corpus in the world.
It was the first text corpus of its size to be made widely available.
It is available from a wide range of different sources.
It is widely regarded as a 'standard reference corpus' for the English language.
It has been licensed to over 1300 institutions throughout the world, over 1800 users have signed on for access to it through the BNCweb online interface, etc.
A failure:
It never reached 100 million words! (98,300,000)
The design criteria were never totally achieved.
It hardly ever contains complete texts.
The spoken materials are poorly transcribed.
The metadata are incomplete and can be erroneous.
The part-of-speech tagging contains many errors.
It is out of date! (dating from the late 20th century)
Leech θդüˤϡtriumph γ˼Ƥ褦ˡӤդ줿ߤʤäƤΥѥԽˤĤơФ褫äФ褫äȤθȤ⤤ȿ¿ƤΤŪǤ롥BNC ΥդѤ줿ץ CLAWS4 ٤97%ۤɤHoffmann et al. 43 ˤȡ98--99%ˤȤΤϡ϶ä٤ȤȻפäƤѥϤ礭ΤǿѡȤΥ顼ȤϤäƤ300ˤΤܤȤ¤ϸȤƤäȤХѥˤĤƤϡѥΤ1ۤɤޤʤäȡǡ transcription μäȡѤǡեޥå TEI äȤФΥդˤɬŬڤǤʤäȡʤɤƤ
ʤǤ⡤ʳ鸽ߤ˻ޤǰӤƤ³Ƥɽ (representativeness) ˤĤơBNC ǤϴṲ̄ʤäȤˡˤޤƤʳ顤ꤹ Text Domain ΥХ䥵˴ؤŤͤƤȤϤ褯ΤƤ롥1桼ȤƤϡ¤줿ΤʤǡɽݤȤϰζȤɾƤ뤬Leech ˤȤäƤϡǤ¤ΤȤϤäȤȿ̤Ȥơ̤ۤʤäȤפ褦Ʊˡ䤫ʸĴǤϤäBNC Ӥ¾Τ٤Ƥ絬ϥѥɽۤɽŻ뤷ƤʤȽƤ༫ȤҤ٤Ƥ褦ˡѥɽˤĤȼϤäƤ뤬ǽŪˤ "impressionistic" ȽǤȹͤƤ褦ǤꡤˤޤƤˤ衤Leech ɽؤμǰζˡ٤ʥץեåʥꥺ
ʤ[2012-07-05-1]ε#1165. ѹǥѥ椬ˤʤäطʡפǿ줿̤ꡤǰʤBNC³ԤϤʤȤȤLeech Ƥ
礭ۤʤ뤬Ѹ쥳ѥ The LAEME Corpus ɽˤĤơ[2012-10-10-1], [2012-10-11-1]εǹͻΤǡȤ
Hoffmann, Sebastian, Stefan Evert, Nicholas Smith, David Lee, and Ylva Berglund Prytz. Corpus Linguistics with BNCweb : A Practical Guide. Frankfurt am Main: Peter Lang, 2008.
[2012-10-10-1], [2012-10-11-1]εǡThe LAEME Corpus ɽˤĤƼꤢɾȤƤϡСƤȻȤߤɽ»ʤƤΤΡѤǤѸ쥳ѥȤƤηŪԤޤ줿絬ϤΥѥǤꡤʬդʧäǸ츦˳Ѥ٤ġǤ롥The LAEME Corpus β٤Ϥ뤷¾Υѥˤ䴰ܻؤ٤ȤϹͤ뤬Ū˸椹ݤɬŪˤĤޤȤ³θɾʤȥեǤ롥
˸ؤϡβξ֤ѻȤ˲ݤƤ롥ȤˤϡߤȤˤϸʤ³ĤޤȤMilroy (45) λŦ˸ظ2Ĥθ³ (limitations of historical inquiry)
[P]ast states of language are attested in writing, rather than in speech . . . [W]ritten language tends to be message-oriented and is deprived of the social and situational contexts in which speech events occur.
[H]istorical data have been accidentally preserved and are therefore not equally representative of all aspects of the language of past states . . . . Some styles and varieties may therefore be over-represented in the data, while others are under-represented . . . . For some periods of time there may be a great deal of surviving information: for other periods there may be very little or none at all.
ۤ³ǤϤ뤬Ϥ뤤ϹˤǤŤϤϡˡǤʤƤ롥ΤʤǤ⡤Smith Ϥο (1) դäդδط뤳ȡ(2) ̻ˤȳ̻ˤбܤ뤳ȡ(3) ߤθβؤαѤβǽõ뤳ȡνŦƤ롥
Ȥ櫓 (3) ˤĤƤϡǯҲؤˤѲ®˿ʤߡθβؤαѤˤʤ褦ˤʤäƤLabov ʸɸ "On the Use of the Present to Explain the Past" ˡľ٣ʪäƤ롥
ȴϢˡǤ uniformitarian_principle ưθ§ˤ̤˲Ф˱ѸʸDenison et al. ԽΤȤˡǯǤ줿ȤդäƤ
Milroy, James. Linguistic Variation and Change: On the Historical Sociolinguistics of English. Oxford: Blackwell, 1992.
Smith, Jeremy J. An Historical Study of English: Function, Form and Change. London: Routledge, 1996.
Labov, William. "On the Use of the Present to Explain the Past." Readings in Historical Phonology: Chapters in the Theory of Sound Change. Ed. Philip Baldi and Ronald N. Werth. Philadelphia: U of Pennsylvania P, 1978. 275--312.
Denison, David, Ricardo Bermúdez-Otero, Chris McCully, and Emma Moore, eds. Analysing Older English. Cambridge: CUP, 2012.
ε[2012-10-10-1]˰³The LAEME Corpus ɽꡥϡΤˤƱѥʸˡͿƤ (tagged words) οˤꡤ头Ȥɽͤ롥ޤɽǤ褦
Table 2: Dialectal and Diachronic Distribution of Linguistic Evidence by Number of Tagged Words
| C12b | C13a | C13b | C14a | Total | |
|---|---|---|---|---|---|
| N | 0 (0.000%) | 362 (0.062) | 0 (0.000) | 52,883 (9.083) | 53,245 (9.146) |
| NEM | 11,342 (1.948) | 0 (0.000) | 3,980 (0.684) | 2,344 (0.403) | 17,666 (3.034) |
| NWM | 0 (0.000) | 58,332 (10.019) | 16,173 (2.778) | 0 (0.000) | 74,505 (12.797) |
| SEM | 40,082 (6.885) | 26,722 (4.590) | 21,921 (3.765) | 31,408 (5.395) | 120,133 (20.634) |
| SWM | 1,030 (0.177) | 90,400 (15.527) | 106,981 (18.375) | 108 (0.019) | 198,519 (34.098) |
| SW | 1,168 (0.201) | 2,610 (0.448) | 46,032 (7.907) | 30,517 (5.242) | 80,327 (13.797) |
| SE | 0 (0.000) | 4,043 (0.694) | 3,199 (0.549) | 30,561 (5.249) | 37,803 (6.493) |
| Total | 53,622 (9.210) | 182,469 (31.341) | 198,286 (34.058) | 147,821 (25.390) | 582,198 (100.000) |

δؿ濴ϽѸηǤ롥λ˴ؿļԤˤȤäƤϡLAEME ԼԤˤСȯ /ˈleɪmiː/ ˤȤ The LAEME Corpus (Text Database) оϡƱ˴ؤ븦ĶġȤơ¤˴ޤ롥LAEME ˤĤƤϡܥ֥Ǥ laeme εǺΤꤢƤȤ櫓ġȤƤβǽõꡤĥ٤#846. HelMapperUK --- hellog ͤαѹϿ CGI ([2011-08-21-1]) #856. LAEME text database Υǡȥƥȵϡ ([2011-08-31-1]) #942. LAEME Index of Sources θġ ([2011-11-25-1]) #1057. LAEME Index of Sources θġ Ver. 2 ([2012-03-19-1]) ɽƤ
繩ˤȤäƻμ줬ʤ褦ˡԤˤȤäƥġθǤ롥Ū The LAEME Corpus ȤäƤ뤦ˡΤȤפȤɤΤ褦ʥѥʤΤΤꤿʤäƤ[2010-11-16-1]ε#568. ѥȱѸ쥳ѥפǼ̤ꡤѥμ礿ħ1Ĥ representativeness ɽˤ롥ϡѥɾΤλɸ1ĤǤ⤢롥˥ѥˤɽγݤˤĤƤϡ#531. OED ΰѥǡѥȤƻȤ뤫 ([2010-10-10-1]) #1243. ٤θ̻ŪΤˡ ([2012-09-21-1]) ǤƤǤ The LAEME Corpus Ƥ롥СƤʬۤˤĤƤϡ#856. LAEME text database Υǡȥƥȵϡ ([2011-08-31-1]) ǺΤꤢʬ˲äƻʬޤʤ The LAEME Corpus ΥġʬϤߤ
ޤϡϿƤƥȤοͤ롥ѥ "scribal text" Ȥñ̤ǥƥȤϿƤ뤬Ȼˤäʬ̤ȡФ礬狼롥ʤʬȻʬϤ켫ΤˡʤΤʲǤϡŪʶʬʤȤϤäƤ⤢٤κϤ뤬ˤȤơ7Ĥء4ĤؤʬƤ롥ʤ N (Northern), NEM (North-East Midland), NWM (North-West Midland), SEM (South-East Midland), SWM (South-West Midland), SW (Southwestern), SE (Southeastern) ء C12b 12ȾˡC13a, C13b, C14a ءѸʬˤĤƤϡ#130. Ѹʬ ([2009-09-04-1]) ⻲ȡ
Table 1: Dialectal and Diachronic Distribution of Linguistic Evidence by Number of Texts
| C12b | C13a | C13b | C14a | Total | |
|---|---|---|---|---|---|
| N | 0 (0.00%) | 1 (0.86) | 0 (0.00) | 7 (6.03) | 8 (6.90) |
| NEM | 1 (0.86) | 0 (0.00) | 5 (4.31) | 2 (1.72) | 8 (6.90) |
| NWM | 0 (0.00) | 9 (7.76) | 5 (4.31) | 0 (0.00) | 14 (12.07) |
| SEM | 4 (3.45) | 7 (6.03) | 14 (12.07) | 7 (6.03) | 32 (27.59) |
| SWM | 2 (1.72) | 13 (11.21) | 17 (14.66) | 1 (0.86) | 33 (28.45) |
| SW | 3 (2.59) | 5 (4.31) | 7 (6.03) | 2 (1.72) | 17 (14.66) |
| SE | 0 (0.00) | 2 (1.72) | 1 (0.86) | 1 (0.86) | 4 (3.45) |
| Total | 10 (8.62) | 37 (31.90) | 49 (42.24) | 20 (17.24) | 116 (100.00) |
ε#1242. -ate ưζܹԡ ([2012-09-20-1]) #1239. Frequency Actuation Hypothesis ([2012-09-17-1]) Ǽ夲 Phillips θΤ褦ˡ٤θѲθˤ¿ʴؿƤ뤬ˡѤʵȤơ٤켫Τ̻ŪѤȤ¤ɤΤ褦˹ͤФ褤ΤȤ꤬롥θǤϤʤäΤ뤤ʬͤˤϡ1--2λ礭ǤϤʤľ롥1λǤϤɤ2ǤϤɤȹͤȡɤޤľΤϤʤϤʤPhillips (225--26) ϡˤĤƼΤ褦˳ڴѤƤ롥
The words' frequencies are based on present-day English, but the general pattern of relative frequencies probably holds for the English in our data base (1755--1993) as well. For example, I would be very surprised if the 3-syllable verbs with CELEX frequencies over 100 --- concentrate, demonstrate, illustrate, contemplate, compensate, designate, and alternate --- were not also much more common in 1755 than those with frequencies of 0 --- altercate, auscultate, condensate, defalcate, eructate, exculpate, expuergate, extirpate, fecundate, etc.
2;λˤƤʤ顤٤100ʾθ0θ٤ȤΤ绨Ĥˤ褦˻פ롥ΤˡPhillips ϼºݤʬϤǤ101ʾ塤10--1001--10ȤӤʬѤƤꡤ绨Ĥپ绨ĤʤޤޤѤ뿵ŤϼƤ롥⤷ä10--100դ٥٥θܺ٤Ĵ٤褦ȤΤǤС2δ֤ˤʤ٤ѲƤǽϤ롥Phillips ʤ餺Ȥ⡤٤Ѥ̻Ū˴ؿï⤬ͤϤ
˻פĤñʲƤϡƻɽǤ礭ʥѥѤɽ뤳ȤǤ롥ƤȤƤñºݤ˿ԤΤϰ֤֤⤫롥ֻٸꤷѸǤСѥѰդɽμưǤѸǤֻ variation 椨 lemmatise Ƥʤ¤ϸФñ̤ǤɽҤޤ夬ŤʤФʤۤɡѥ˴ޤޤƥȤ representativeness Ͽˤʤ롥ӤäݤɽǤ⡤ʤϤۤ褤ƤߤȻפäƤ롥뤤ϡˤäƤϤǤˤ
ʤѤˤ CELEX Ȥñǡ١ϡѸθ֤˴ؤŪʸǤ褯ȤƤΤǤ롥ܺ٤ϡCELEX2 ȡޤ٤̻֤δطˤĤƤϡ[2012-05-03-1]ε#1102. Zipf's law ȸοաפȡ
Phillips, Betty S. "Word Frequency and Lexical Diffusion in English Stress Shifts." Germanic Linguistics. Ed. Richard Hogg and Linda van Bergen. Amsterdam: John Benjamins, 1998. 223--32.
ܥ֥Ǥⲿ٤夲Ƥ2Ĥ˱Ѹ쥳ѥ PPCMBE ( Penn Parsed Corpus of Modern British English; see [2010-03-03-1]. ) COHA ( Corpus of Historical American English; see [2010-09-19-1]. ) ˤĤơܻرѸ쥳ѥ٤κǿ˸ΡȤȯɽƤ롥ξԤȤ2010ǯ˸줿ѸΥѥ줾ѼǤ뤳ȡޤԻŪۤʤ뤳Ȥ٤ӤоݤˤŬʤɽϤȤ륳ѥΰŪħ٤뤳Ȥϰ̣
PPCMBE 1700--1914ǯΥꥹѸƥ949,000ǹƤꡤParsed Corpora of Historical English 1ʤƱͤ˹ʸϤ줿Ťб륳ѥȤ³ռǤ롥ͭǥǡꤹɬפ롥COHA 1810--2009ǯΥꥫѸƥ4Ͽ祳ѥǤ롥ϡʸϤϤƤʤCOHA ̵ǥ饤Ǥ뤿Ȥ䤹եꤵƤΤǽʥǡǤʤȤ롥
ѥεϤȤط뤬PPCMBE ɽ (representativeness) 롥PPCMBE ΥѥƥȤ18غ٤ʬषƥǯ10ǯߤǤȤȡȤʤޥܤ¿롥ϡʬ٤ͭյʬϷ̤ФʤȤȤǤꡤѤ˺ݤդפ롥
COHA ΥѥƥȤ Fiction, Popular Magazines, Newspapers, Non-Fiction Books 4绨Ĥ˶ʬƤ롥٤ʬθˤѤǤʤ10ǯߤǤƥޥܤŬڤʥΥƥȤۤƤꡤɽϤ褯ݤƤ롥Fiction ιΨɤλ50%ƤꡤFiction θħä˸áˤѥΤθħ˱ƶͿƤȹͤ졤ʬϤκݤˤϤդפ롥
ܻϡξѥΰʾħѸˤƻӵ顦ǾˤäƼƤ롥CONCE (Corpus of Nineteenth-Century English) Ѥ Kytö and Romaine ԸˤС19δ֡ӵαФγϡ30ǯߤƬ57.1%67.8%ؤäƤȤƱͤĴ COHA PPCMBE 10ǯߤ˻ܤȤԤǤ1810ǯ64.7%1910ǯ74.3%¤äƤ뤳ȤΤ줿ԤǤ1810ǯ79.4%1910ǯ78.0%ޤɤ줬㤷äȤܡp. 56ˡCONCEƱͤ30ǯߤʬϤľȡPPCMBE ǤͭդѲܴۤѻǤۤɤη̤ǤȤ
ѥϤ줾ȼħäƤ롥褯İѤɬפ뤳ȤǧϢơ[2010-06-04-1]εή˵դäƤӵˡפȡ
2ĤλŦѥ---ɽסرѸ쥳ѥ18桤Ѹ쥳ѥز2011ǯ49--59ǡ
Kytö, M. and S. Romaine. "Adjective Comparison in Nineteenth-Century English." Nineteenth-Century English: Stability and Change. Ed. M. Kytö, M. Rydén, and E. Smitterberg. Cambridge: CUP, 2006. 194--214.
츦ˤ corpus ֥ѥפ͡Ƥ뤬McEnery et al. ʷǤ롥
. . . a corpus is a collection of (1) machine-readable (2) authentic texts (including transcripts of spoken data) which is (3) sampled to be (4) representative of a particular language or language variety.
(1) (2) ˤĤƤϤ褽Դ֤˥뤬(3) (4) ˤĤƤϲä "sampled" 뤤 "representative" ȤߤʤˤĤ͡ʰո롥ڤˤƤ뤳ȤǤ
ڤ˱Ѹ쥳ѥˤϡ饤ΤΤǤ롥ʲϡϿɬפʤΤ⤢뤬˥饤ǴؤѤǤѸ쥳ѥ
British National Corpus ʤĤΥեƤ
* BNC ( The British National Corpus )
* BNCweb ̵Ͽ
* BYU-BNC ̵Ͽ
BYU Corpora Brigham Young University, Mark Davies Τ¾Υ饤ѥ
* COCA ( Corpus of Contemporary American English ) ̵Ͽ
* COHA ( Corpus of Historical American English ) ̵Ͽ
* TIME Magazine Corpus of American English ̵Ͽ
Cobuild Concordance and Collocations Sampler
¾ܥ֥ǤϥѥطεȷǺܤƤΤǡͤˤ줿
hellog Υѥν: [2010-09-15-1]
hellog ΥѥϢ: corpus
hellog BNC Ϣ: bnc
McEnery, Tony, Richard Xiao, and Yukio Tono. Corpus-Based Language Studies: An Advanced Resource Book. London: Routledge, 2006.
OED (2nd ed. CD-ROM) ˱Ѹ쥳ѥȤѤȤȯۤäŻǤǤƤ鹭ͭƤºݤ¿θ OED ѥȤƳѤƤ롥⤽⤬ѥȤԤޤ줿櫓ǤϤʤ OED νѥȤߤʤƸ椹뤳Ȥϡɤ줯餤ʤΤƻˤĤΤ뤳Ȥϸ漫ȤƱ餤פȻפΤǡΥơޤ˴Ϣ Hoffmann ʸޤȤƤߤʻ伫ȤƻȤƤ OED ħ褯˸˻ȤäƤ餤ΤǡʬΤ˺ϿȤĤǤսν줿ʸͤˤƤޤ
Hoffmann OED νѥȤѤ뤳ȤǤ뤫ȤФơ4Ĥδ饢ץƤ롥ƴȡб Hoffmann η롥
(1) Selection criteria for the quotations
"a collection of pieces of language that are selected and ordered according to explicit linguistic criteria in order to be used as a sample of the language" (19; cited from Sinclair) Ȥ̩ʥѥ˾Ȥ餻СOED νѥȸʤȤϤǤʤΤˡġθФ첼ǼƤ㷲θФܤŬڤʥѥˤʤʤȤȤϸθü٤η֤̣åפ뷹뤫Ǥ롥äˤ븫ФܤΤǤʤСΤȤ OED ϳƻαѸɽƤȹͤ졤ѥȤƳѤ뤳ȤǤ롥
(2) Representativeness and balance of the quotations
OED ϼºݤ˲餫ŵƤ "true quotations" (20) Ǥ롥ԼԤˤäƺ줿ʤǤϤʤϤƾʤޤŵΥ¿ˤ錄ꡤüʸغʤ˸¤ʤɤиʤΤǡ˴ؤƤ "representative" ȸäƤ褤ƥ뤬츦ˤȤäŬڤʳʬۤƤ櫓ǤϤʤΤǡ"balanced" Ȥϸʤ㤨 Shakespeare 1ͤ33,000Ƥʤɤ롥OED ѥȤƸΩƤˤϡ"balance" դפ롥
(3) Reliability of the data format
ʸΰάƤ褦㤬ʿѤ20?25%ۤɤ롥ۤȤɤξάǤʸι¤ƤʤˤŬڤʾάʸι¤ѲƤޤäƤʸ⤢롥ʾι¤Ĵ٤뤿 OED ѤˤϡդɬפǤ롥
(4) Quantification of the results
1ǯդ˥ץåȤȡ174000ۤ뾮ԡ1910000ۤԡǧ뤬20ˤϷ㸺롥ǡοϻˤ餺13٤Ȱǡ20㤬ĹʤΤܤαޤ٤Ǥ롥240ۤʽǤ180ۤɤäˤȤȾ嵭ʿѸơOED ˴ޤޤ3300?3500ȿꤵ롥OED ѥȤѤˤϡ19ä¿ȤʤɤդƸ̤᤹٤
Ǹ Hoffmann ηѤ (26) OED νϸѲη绨ĤŪɽ魯ѥȤƸѲˤȤäͭѤǤ롤ȤQŪʷŪʿФƤƻͤˤʤä
Although the OED quotations database is not a completely balanced and representative corpus, it can nevertheless provide the linguist with a wealth of useful information. The data it contains chiefly represents naturally occurring language, and the time-span covered is unmatched by any other source of computerized data. Even though over 20 per cent of all its quotations have been shortened, the large majority of these deletions is unlikely to distort the results of many diachronic studies of linguistic features. Given the nature of the data, normalized frequency counts might suggest an inappropriate level of precision, but tendencies in the development over time can nevertheless be expressed in quantitative terms. (26)
The Oxford English Dictionary. 2nd ed. CD-ROM. Version 3.1. Oxford: OUP, 2004.
Hoffmann, Sebastian. "Using the OED quotations database as a Corpus --- A Linguistic Appraisal." ICAME Journal 28 (April 2004): 17--30. Available online at http://icame.uib.no/ij28/index.html .
Tanabe, Harumi. "The Rivalry of give up and its Synonymous Verbs in Modern English." Language Change and Variation from Old English and Late Modern English: A Festschrift for Minoji Akimoto. Ed. Merja Kytö, John Scahill, and Harumi Tanabe. Bern: Peter Lang, 2010. 253--75.
Powered by WinChalow1.0rc4 based on chalow