ѥˤȤäɽ (representativeness) ̿Ǥ뤳Ȥϡѥ ([2010-11-16-1]) 餫Ǥ뤷ε#1279. BNC ζߤȼߡ ([2012-10-28-1]) ǾҲ𤷤 Leech Ȥ櫓ĥƤǤ롥McEnery et al. (13) ϡɽˤĤơLeech ͤˤʤ "a corpus is thought to be representative of the language variety it is supposed to represent if the findings based on its contents can be generalized to the said language variety" ȽҤ٤Ƥ롥
ɽŪ˹ͤƤߤ褦㤨 BNC åȤȤ褦ʡ奤ꥹѸȤŪѼϿ륳ѥ (general corpus) ɽϤɤΤ褦ˤΤäդȽդγͤȡ줾50%Ĥ˳꿶뤳Ȥϡ奤ꥹѸɽ«ƤLeech ɽǤ "impressionistic" Ȥʤ餶ʤνִ֤˹ԤʤƤ븽奤ꥹѸΰŪʬäդˤƤǤϤʤ⤷ȤСäեѥγ㤨80%ۤɤꤹۤɽݤǤΤǤϤʤΤȤʤ븽奤ꥹѸľܤĤळȤǤʤʾ塤ɽεϹԤͤޤäƤޤ
ѥä˰̥ѥˤɽȤˡ balance sampling Ȥ2Ĥγǰʬƹͤ뤳Ȥ롥McEnery et al. (13) Ǥϡ"the representativeness of most corpora is to a great extent determined by two factors: the range of genres included in a corpus (i.e. balance . . .) and how the text chunks for each genre are selected (i.e. sampling . . .)" Ƥ롥
balance ȤϡBNC ѸǤȤ domain genre Ȥʬ˴ؤΤǤ롥㤨С奤ꥹѸΥѥɸ֤ʤ⡤ꥹοʹαѸѥϡrepresentativeness 롥奤ꥹѸˤϽդǤʤäդ⤢뤷ԤˤĤƤϿʹѸǤʤʸرѸ⤢ŻҥѸ⤢뤷㤤ʪѸ⤢СѸ⤢롥Τ domain genre θ줿Ȼפ̤ƤĤ text domain ΤʹѸ˸¤äƤ⡤֥ɤ⤢й⤢롥1ĤοʹǤ⡤Ҳ̡ݡ̡ʤɤ̤ɬפϤʤΤҲ̤Ǥйȹݵζ̤ϤɤŪˤϤɤޤǤʬ롥äդǤƱͤ˺ʬ䤷ʤƤСĿ (idiolect) ˸Ŀˤ register ̤θ졤ʤɤΥȥؤȽ夷ƤޤºݤΥѥϡQŪʥ٥Ŷ뤳Ȥˤʤ뤬־QŪפ "impressionistic" ϤۤƱ
sampling Ȥɽ뤿μˡǤ롥ΤθŪħƸ褦ˡ̤ˤƹθäʤ顤ѥ˳ domain ۤ뤿ȼǤ롥ˤϡsampling unit ȤƲꤹ뤫ŵŪˤϡܡʹʤɤʤȤƤñ̡ˡΤ褦ñ̤ꥹȲȤϰ (sampling frame) ɤޤǤꤹ뤫ǯؤθ䡤٥ȥ顼ܤؤθʤɡˡɸܼϴʥˤ뤫٤ηϲäǤΥˤ뤫ɤۤ뤫ʤɤΡŪŪ꤬ޤޤ롥
ɽ˴ؤ⤦1ĤγǰȤơclosure 뤤 saturation ȸƤФΤ⤢롥McEnery et al. (16) ˤС"Closure/saturation for a particular linguistic feature (e.g. size of lexicon) of a variety of language (e.g. computer manuals) means that the feature appears to be finite or is subject to very limited variation beyond a certain point." Ƥ롥ʿСʾ女ѥεϤ礭Ƥ⡤ùγѤʤȤϤãСΥѥ saturated Ǥȹͤ롥ɽλɸȤƤϡbalance saturation ΤۤƤȤŦ⤢뤬saturation ϼȤƸäǰƬˤꡤ¾θܤؤαѤϻߤƤʤΤǤ롥
ɽϡ女ѥ̿ǤȤϤäƤ⡤ԤȤ餤Ϥ롥ݤ뤿ʤˡʤ٤ƤΥѥԻԤΩϤƬˤѥϼԻƤ롥Ṳ̄ˤơҤԻȻѤ³Ƥ椭Υϥ٤ʳˤΤ⤷ʤ
McEnery, Tony, Richard Xiao, and Yukio Tono. Corpus-Based Language Studies: An Advanced Resource Book. London: Routledge, 2006.
108--114֤ˤ錄ꡤΩرѸ춵鸦ˤŤǡLancaster ̾ Geoffrey Leech θֱ줿ϡ2ܤ "The British National Corpus: Both a Triumph and a Failure" ꤹֱΤߤλääİ˹ԤäBNC ԼԤκäʤɡ⤷ää
̾ˤ "triumph" "failure" ˤĤơLeech Ϥ줾켡Τ褦ʹܤƤ
A triumph:
It has been claimed that the BNC is the most widely used corpus in the world.
It was the first text corpus of its size to be made widely available.
It is available from a wide range of different sources.
It is widely regarded as a 'standard reference corpus' for the English language.
It has been licensed to over 1300 institutions throughout the world, over 1800 users have signed on for access to it through the BNCweb online interface, etc.
A failure:
It never reached 100 million words! (98,300,000)
The design criteria were never totally achieved.
It hardly ever contains complete texts.
The spoken materials are poorly transcribed.
The metadata are incomplete and can be erroneous.
The part-of-speech tagging contains many errors.
It is out of date! (dating from the late 20th century)
Leech θդüˤϡtriumph γ˼Ƥ褦ˡӤդ줿ߤʤäƤΥѥԽˤĤơФ褫äФ褫äȤθȤ⤤ȿ¿ƤΤŪǤ롥BNC ΥդѤ줿ץ CLAWS4 ٤97%ۤɤHoffmann et al. 43 ˤȡ98--99%ˤȤΤϡ϶ä٤ȤȻפäƤѥϤ礭ΤǿѡȤΥ顼ȤϤäƤ300ˤΤܤȤ¤ϸȤƤäȤХѥˤĤƤϡѥΤ1ۤɤޤʤäȡǡ transcription μäȡѤǡեޥå TEI äȤФΥդˤɬŬڤǤʤäȡʤɤƤ
ʤǤ⡤ʳ鸽ߤ˻ޤǰӤƤ³Ƥɽ (representativeness) ˤĤơBNC ǤϴṲ̄ʤäȤˡˤޤƤʳ顤ꤹ Text Domain ΥХ䥵˴ؤŤͤƤȤϤ褯ΤƤ롥1桼ȤƤϡ¤줿ΤʤǡɽݤȤϰζȤɾƤ뤬Leech ˤȤäƤϡǤ¤ΤȤϤäȤȿ̤Ȥơ̤ۤʤäȤפ褦Ʊˡ䤫ʸĴǤϤäBNC Ӥ¾Τ٤Ƥ絬ϥѥɽۤɽŻ뤷ƤʤȽƤ༫ȤҤ٤Ƥ褦ˡѥɽˤĤȼϤäƤ뤬ǽŪˤ "impressionistic" ȽǤȹͤƤ褦ǤꡤˤޤƤˤ衤Leech ɽؤμǰζˡ٤ʥץեåʥꥺ
ʤ[2012-07-05-1]ε#1165. ѹǥѥ椬ˤʤäطʡפǿ줿̤ꡤǰʤBNC³ԤϤʤȤȤLeech Ƥ
礭ۤʤ뤬Ѹ쥳ѥ The LAEME Corpus ɽˤĤơ[2012-10-10-1], [2012-10-11-1]εǹͻΤǡȤ
Hoffmann, Sebastian, Stefan Evert, Nicholas Smith, David Lee, and Ylva Berglund Prytz. Corpus Linguistics with BNCweb : A Practical Guide. Frankfurt am Main: Peter Lang, 2008.
ѥؤߤޤʤʬʬˡϢϥ־뤳Ȥ¿ʤ褦ˤפ뤬դ˾¿ơȽǤ˺롥ƼʬΤǤʥޤȤƤȻפΤسΥԡɤˤĤƹԤʤ䤬Ǥ褯Ѥ BNC ˴ϢΤ濴ˡŪǤϤ뤬ĥ롥ޤȤϫϡŤ뼰ˤɤ뤫ɸΤۤΨŪȤˤʤĤĤ롦
1. BNC ե
BNCweb ̵Ͽ
BYU-BNC ̵Ͽ
BNC ( The British National Corpus )
2. BNC Υե
Quick Reference for Simple Query Syntax (PDF)
Reference Guide for the British National Corpus (XML Edition)
Reference Guide ܼ
* 6.5 Guidelines to the Wordclass Tagging
* The BNC Basic (C5) Tagset
* 9.8 Simplified Wordclass Tags
* 9.7 Contracted forms and multiwords
* 1 Design of the Corpus
* 9.6 Text and genre classification code
3. ѥϢ祵
David Lee ˤ Bookmarks for Corpus-based Linguists
* Corpora, Collections, Data Archives
* Software, Tools, Frequency Lists, etc.
* References, Papers, Journals
* Conferences & Project
4. hellog ε
#568. ѥȱѸ쥳ѥ: [2010-11-16-1]
#506. CoRD --- Ѹ˥ѥξ: [2010-09-15-1]
#308. Ѹκѱñꥹȡ: [2010-03-01-1]
ѥϢ: corpus
BNC Ϣ: bnc
COCA Ϣ: coca
5. ġ
Corpus Frequency Wizard
Paul Rayson's Log-likelihood Calculator
VassarStats
hellog Ρ#711. Log-Likelihood Tester CGI, Ver. 2: [2011-04-08-1]
Hoffmann, Sebastian, Stefan Evert, Nicholas Smith, David Lee, and Ylva Berglund Prytz. Corpus Linguistics with BNCweb : A Practical Guide. Frankfurt am Main: Peter Lang, 2008.
ɸΤ褦 here there 1ǤȤֻ2ǤȤʣ¿롥ϡhere this ȡthere it that ɤؤơֻθ˲Ọ̇̄ŪɸθϤ줾 by this, of this, to that, with that ۤɤ̣롥Ǥ˷ĥä뤬űѸ줫ѸˤƤϤ褯Ѥ졤μ٤ϤषƤۤɤǤ롥17ʹߤϵ˸äƤ椭Τ褦ʸ¤줿Ѱ (register) ؤɤޤ줿ͳȤƤϡѸι¤ȤŵŪǤʤȤĤޤ礫ʬϤؤαѸμήȿȤŦƤ (Rissanen 127) ʸˡȤơޤǸꤵ줿֤ǼѤ줿ϡtherefore ΤߤȤäƤ褤
ѸdzǧѰФϡǤѸˤ˨꤬롥here-, there- ʣϡѸǤϤޤ̤˻ȤƤ뤬ǤߤˡΧʸǤλѤݤäƤ롥ʲϡRissanen (127) Helsinki Corpus ˤĴ̤Ǥʿ١åο1٤ɽ魯ˡ
| Statutes | Other texts | |
|---|---|---|
| ME4 (1420--1500) | 68 (60) | 621 (31) |
| EModE1 (1500--70) | 77 (65) | 503 (28) |
| EModE2 (1570--1640) | 84 (71) | 461 (26) |
| EModE3 (1640--1710) | 126 (96) | 191 (12) |
[2012-10-10-1], [2012-10-11-1]εǡThe LAEME Corpus ɽˤĤƼꤢɾȤƤϡСƤȻȤߤɽ»ʤƤΤΡѤǤѸ쥳ѥȤƤηŪԤޤ줿絬ϤΥѥǤꡤʬդʧäǸ츦˳Ѥ٤ġǤ롥The LAEME Corpus β٤Ϥ뤷¾Υѥˤ䴰ܻؤ٤ȤϹͤ뤬Ū˸椹ݤɬŪˤĤޤȤ³θɾʤȥեǤ롥
˸ؤϡβξ֤ѻȤ˲ݤƤ롥ȤˤϡߤȤˤϸʤ³ĤޤȤMilroy (45) λŦ˸ظ2Ĥθ³ (limitations of historical inquiry)
[P]ast states of language are attested in writing, rather than in speech . . . [W]ritten language tends to be message-oriented and is deprived of the social and situational contexts in which speech events occur.
[H]istorical data have been accidentally preserved and are therefore not equally representative of all aspects of the language of past states . . . . Some styles and varieties may therefore be over-represented in the data, while others are under-represented . . . . For some periods of time there may be a great deal of surviving information: for other periods there may be very little or none at all.
ۤ³ǤϤ뤬Ϥ뤤ϹˤǤŤϤϡˡǤʤƤ롥ΤʤǤ⡤Smith Ϥο (1) դäդδط뤳ȡ(2) ̻ˤȳ̻ˤбܤ뤳ȡ(3) ߤθβؤαѤβǽõ뤳ȡνŦƤ롥
Ȥ櫓 (3) ˤĤƤϡǯҲؤˤѲ®˿ʤߡθβؤαѤˤʤ褦ˤʤäƤLabov ʸɸ "On the Use of the Present to Explain the Past" ˡľ٣ʪäƤ롥
ȴϢˡǤ uniformitarian_principle ưθ§ˤ̤˲Ф˱ѸʸDenison et al. ԽΤȤˡǯǤ줿ȤդäƤ
Milroy, James. Linguistic Variation and Change: On the Historical Sociolinguistics of English. Oxford: Blackwell, 1992.
Smith, Jeremy J. An Historical Study of English: Function, Form and Change. London: Routledge, 1996.
Labov, William. "On the Use of the Present to Explain the Past." Readings in Historical Phonology: Chapters in the Theory of Sound Change. Ed. Philip Baldi and Ronald N. Werth. Philadelphia: U of Pennsylvania P, 1978. 275--312.
Denison, David, Ricardo Bermúdez-Otero, Chris McCully, and Emma Moore, eds. Analysing Older English. Cambridge: CUP, 2012.
ε[2012-10-10-1]˰³The LAEME Corpus ɽꡥϡΤˤƱѥʸˡͿƤ (tagged words) οˤꡤ头Ȥɽͤ롥ޤɽǤ褦
Table 2: Dialectal and Diachronic Distribution of Linguistic Evidence by Number of Tagged Words
| C12b | C13a | C13b | C14a | Total | |
|---|---|---|---|---|---|
| N | 0 (0.000%) | 362 (0.062) | 0 (0.000) | 52,883 (9.083) | 53,245 (9.146) |
| NEM | 11,342 (1.948) | 0 (0.000) | 3,980 (0.684) | 2,344 (0.403) | 17,666 (3.034) |
| NWM | 0 (0.000) | 58,332 (10.019) | 16,173 (2.778) | 0 (0.000) | 74,505 (12.797) |
| SEM | 40,082 (6.885) | 26,722 (4.590) | 21,921 (3.765) | 31,408 (5.395) | 120,133 (20.634) |
| SWM | 1,030 (0.177) | 90,400 (15.527) | 106,981 (18.375) | 108 (0.019) | 198,519 (34.098) |
| SW | 1,168 (0.201) | 2,610 (0.448) | 46,032 (7.907) | 30,517 (5.242) | 80,327 (13.797) |
| SE | 0 (0.000) | 4,043 (0.694) | 3,199 (0.549) | 30,561 (5.249) | 37,803 (6.493) |
| Total | 53,622 (9.210) | 182,469 (31.341) | 198,286 (34.058) | 147,821 (25.390) | 582,198 (100.000) |

δؿ濴ϽѸηǤ롥λ˴ؿļԤˤȤäƤϡLAEME ԼԤˤСȯ /ˈleɪmiː/ ˤȤ The LAEME Corpus (Text Database) оϡƱ˴ؤ븦ĶġȤơ¤˴ޤ롥LAEME ˤĤƤϡܥ֥Ǥ laeme εǺΤꤢƤȤ櫓ġȤƤβǽõꡤĥ٤#846. HelMapperUK --- hellog ͤαѹϿ CGI ([2011-08-21-1]) #856. LAEME text database Υǡȥƥȵϡ ([2011-08-31-1]) #942. LAEME Index of Sources θġ ([2011-11-25-1]) #1057. LAEME Index of Sources θġ Ver. 2 ([2012-03-19-1]) ɽƤ
繩ˤȤäƻμ줬ʤ褦ˡԤˤȤäƥġθǤ롥Ū The LAEME Corpus ȤäƤ뤦ˡΤȤפȤɤΤ褦ʥѥʤΤΤꤿʤäƤ[2010-11-16-1]ε#568. ѥȱѸ쥳ѥפǼ̤ꡤѥμ礿ħ1Ĥ representativeness ɽˤ롥ϡѥɾΤλɸ1ĤǤ⤢롥˥ѥˤɽγݤˤĤƤϡ#531. OED ΰѥǡѥȤƻȤ뤫 ([2010-10-10-1]) #1243. ٤θ̻ŪΤˡ ([2012-09-21-1]) ǤƤǤ The LAEME Corpus Ƥ롥СƤʬۤˤĤƤϡ#856. LAEME text database Υǡȥƥȵϡ ([2011-08-31-1]) ǺΤꤢʬ˲äƻʬޤʤ The LAEME Corpus ΥġʬϤߤ
ޤϡϿƤƥȤοͤ롥ѥ "scribal text" Ȥñ̤ǥƥȤϿƤ뤬Ȼˤäʬ̤ȡФ礬狼롥ʤʬȻʬϤ켫ΤˡʤΤʲǤϡŪʶʬʤȤϤäƤ⤢٤κϤ뤬ˤȤơ7Ĥء4ĤؤʬƤ롥ʤ N (Northern), NEM (North-East Midland), NWM (North-West Midland), SEM (South-East Midland), SWM (South-West Midland), SW (Southwestern), SE (Southeastern) ء C12b 12ȾˡC13a, C13b, C14a ءѸʬˤĤƤϡ#130. Ѹʬ ([2009-09-04-1]) ⻲ȡ
Table 1: Dialectal and Diachronic Distribution of Linguistic Evidence by Number of Texts
| C12b | C13a | C13b | C14a | Total | |
|---|---|---|---|---|---|
| N | 0 (0.00%) | 1 (0.86) | 0 (0.00) | 7 (6.03) | 8 (6.90) |
| NEM | 1 (0.86) | 0 (0.00) | 5 (4.31) | 2 (1.72) | 8 (6.90) |
| NWM | 0 (0.00) | 9 (7.76) | 5 (4.31) | 0 (0.00) | 14 (12.07) |
| SEM | 4 (3.45) | 7 (6.03) | 14 (12.07) | 7 (6.03) | 32 (27.59) |
| SWM | 2 (1.72) | 13 (11.21) | 17 (14.66) | 1 (0.86) | 33 (28.45) |
| SW | 3 (2.59) | 5 (4.31) | 7 (6.03) | 2 (1.72) | 17 (14.66) |
| SE | 0 (0.00) | 2 (1.72) | 1 (0.86) | 1 (0.86) | 4 (3.45) |
| Total | 10 (8.62) | 37 (31.90) | 49 (42.24) | 20 (17.24) | 116 (100.00) |
ε#1242. -ate ưζܹԡ ([2012-09-20-1]) #1239. Frequency Actuation Hypothesis ([2012-09-17-1]) Ǽ夲 Phillips θΤ褦ˡ٤θѲθˤ¿ʴؿƤ뤬ˡѤʵȤơ٤켫Τ̻ŪѤȤ¤ɤΤ褦˹ͤФ褤ΤȤ꤬롥θǤϤʤäΤ뤤ʬͤˤϡ1--2λ礭ǤϤʤľ롥1λǤϤɤ2ǤϤɤȹͤȡɤޤľΤϤʤϤʤPhillips (225--26) ϡˤĤƼΤ褦˳ڴѤƤ롥
The words' frequencies are based on present-day English, but the general pattern of relative frequencies probably holds for the English in our data base (1755--1993) as well. For example, I would be very surprised if the 3-syllable verbs with CELEX frequencies over 100 --- concentrate, demonstrate, illustrate, contemplate, compensate, designate, and alternate --- were not also much more common in 1755 than those with frequencies of 0 --- altercate, auscultate, condensate, defalcate, eructate, exculpate, expuergate, extirpate, fecundate, etc.
2;λˤƤʤ顤٤100ʾθ0θ٤ȤΤ绨Ĥˤ褦˻פ롥ΤˡPhillips ϼºݤʬϤǤ101ʾ塤10--1001--10ȤӤʬѤƤꡤ绨Ĥپ绨ĤʤޤޤѤ뿵ŤϼƤ롥⤷ä10--100դ٥٥θܺ٤Ĵ٤褦ȤΤǤС2δ֤ˤʤ٤ѲƤǽϤ롥Phillips ʤ餺Ȥ⡤٤Ѥ̻Ū˴ؿï⤬ͤϤ
˻פĤñʲƤϡƻɽǤ礭ʥѥѤɽ뤳ȤǤ롥ƤȤƤñºݤ˿ԤΤϰ֤֤⤫롥ֻٸꤷѸǤСѥѰդɽμưǤѸǤֻ variation 椨 lemmatise Ƥʤ¤ϸФñ̤ǤɽҤޤ夬ŤʤФʤۤɡѥ˴ޤޤƥȤ representativeness Ͽˤʤ롥ӤäݤɽǤ⡤ʤϤۤ褤ƤߤȻפäƤ롥뤤ϡˤäƤϤǤˤ
ʤѤˤ CELEX Ȥñǡ١ϡѸθ֤˴ؤŪʸǤ褯ȤƤΤǤ롥ܺ٤ϡCELEX2 ȡޤ٤̻֤δطˤĤƤϡ[2012-05-03-1]ε#1102. Zipf's law ȸοաפȡ
Phillips, Betty S. "Word Frequency and Lexical Diffusion in English Stress Shifts." Germanic Linguistics. Ed. Richard Hogg and Linda van Bergen. Amsterdam: John Benjamins, 1998. 223--32.
Ѹˤ whom οˤĤƤϡ¿θ椬롥ѸǤ褯Τ줿ѲǤꡤܥ֥Ǥ ##622,624,860,301,737 γƵǿƤĤƤ´ˤ⤳äΤ ([2010-12-26-1])
ǶθȤƤϡIyeiri and Yaguchi 롥ϡMichael Barlow ԻAthelstan ͭƤ The Corpus of Spoken Professional American English (CSPAE) ˴ŤǤ롥CSPAE ϡ1990ǯ祢ꥫѸäեѥǡ(1) White House ǤεԲ(2) The University of North Carolina ζ(3) إƥȰѰιȲġ(4) ɲƥȰѰιȲĤΡ4ĤξʬƤꡤΤȤ200줫롥ޤCLAWS7 ǥդƤ롥ϡwhom ϷĥäʸΡä˽դˤƻѤȤ뤬ǤϡĥääդȤĶǤɤٻȤΤȤ䤤뤳ȤǤ롥
Ĵ̤˽Сspoken professional American English ˤƤϡwhom οǤʤΤΡޤ٤٤Ǥϸ롥whom Ķˤ餫ʷꡤֻľˤƤϺǤ褯ݤƤʤδĶǤ who λϳ̵ǤϤʤˡֻ (preposition_stranding) ˤϤƤ who Ǥ롥ޤwho(m) ֻŪǤϤʤưŪȤƵǽƤˤϡ礭ɤ줬롥
ȤƤ whom ȴطȤƤ whom ٤ȡԤΤۤब㤷ΤˡɮԤ Rohdenburg ˤ "Complexity Principle (transparency principle)" ѤƤ롥ϡ"[i]n the case of more or less explicit grammatical options the more explicit one(s) will tend to be favored in cognitively more complex environments" (cited in Iyeiri and Yaguchi, p. 185) Ȥǡwhom εƤϤȡطޤʸǧξʣǤꡤŪʳɸ᤹롤ȤȤˤʤ롥
ҤΤȤꡤwhom οϸѸθѲȤƼ夲뤳Ȥ¿Τ褦ˤĤơreferences ˻ͻޤȤƤΤϤ꤬ޤäեѥλѤˤؿ襤ϢơThe Michigan Corpus of Academic Spoken English Ȥѥ⻲ȡ
Iyeiri, Yoko and Michiko Yaguchi. "Relative and Interrogative Who/Whom in Contemporary Professional American English." Germanic Languages and Linguistic Universals. Ed. John Ole Askedal, Ian Roberts, Tomonori Matsushita, and Hiroshi Hasegawa. Tokyo: Senshu University, 2009. 177--91.
رѸ쥳ѥ19˷Ǻܤʸǡ1960ǯư衤ѥؤȤ櫓ѹȯŸƤаޤȤƤǤϡthe University of Birmingham, Lancaster University, the University of Nottingham 3ؤѥؤȯŸ˲̤Ƥ䤬ĴƤꡤѹˤ륳ѥθŸ˾ޤǤ⤬Τ褯Ѥ졤˻ͤˤʤä
ʸˤȡѹǥѥ椬ˤʤäطʤˤϡ5ä (68--69)
(1) Ԥˡ絬ϤʸץȤ˻äŪ;͵äʤˡ
(2) ʸˡʳθФƴƤھä
(3) ǼҤѥμŪʱѡä˼Իˤ˴ؿ
(4) 1990ǯˤϡthe Bank of English, the British National Corpus, the London-Lund Corpus ޤࡤ¿εɼʥѥ˥Ǥ
(5) ѼԤȤϢˤꡤѥʬϤġ뤬äex. Micro-Concord, WordSmith, AntConc, BNCwebˡ
ˤĤơBirmingham ǤϡJohn Sinclair ζϤʻƳϤΤȤݤ줿³Ƥ롥collocation, meaning unit, semantic preference, semantic prosody, discourse analysis, pattern grammar, expressions of evaluation, modal-like expressions ʤɤɤȤ륳ѥ椬˿ʤƤ롥
Lancaster ǤϡGeoffrey Leech, Tony McEnery, Andrew Wilson ʤɤˤ륳ѥγȯȸ椬ʤƤThe Brown Family of Corpora κ˴ؤäۤդץ CLAWS BNCweb γȯUCREL (University Centre for Computer Corpus Research in Language) ΩʤɡŪŪ¦̤ǤĹ롥ߤǤϡŪʸΤȤʤ顤ѸʳθؤȴؿĤĤ롥ǡˤ BNC Τ褦ʵץȤ³ԤϴԤǤʤ褦
Nottingham ǤϡRonald Carter, Michael McCarthy äդؤδؿ顤1990ǯƬ CUP ȶƱơCANCODE (the Cambridge and Nottingham Corpus of Discourse in English) Իθ⡤³͡ʥѥƤNottingham ˤ륳ѥħȤƤϡäդȽդˤʸˡκۡmultimodal corpus ԻʤɤεŪʳ춵ؤαѤ롥
1960ǯ˻夲女ѥؤ1970--1980ǯȯŸη̡1990ǯ˼ήʤʬȤƳΩ21ֲפ˻äƤ롥
Anthony, Laurence, Yasunori Nishina, Kaoru Takahashi, and Michael Handford. "Current Trends in Corpus Linguistics: Voices from Britain."رѸ쥳ѥ19桤Ѹ쥳ѥز2012ǯ67--92ǡ
ε#1160. MRC Psychological Database Ƽפв ([2012-06-30-1]) (3) ǡѸäˤʬ̤ơ줾٤ФˤȡоݤȤʤä92767θΤˤ1졤2졤3졤4ϡ줾13.46%35.40%29.91%15.26%Ǥꡤ碌94.03%ã롥Ȥ櫓23碌65.31%Ǥ롥9;Ȥ絬ϤʸäĴ¤ꡤѸä3ʬ22--3ǤȤȤˤʤ롥
##348,349,355 εǤϡBNC COLT ΥѥѤơǤ٤ι⤤ɴ줫оݤ˲ĴԤʤäĴоݤȤʤäεϤϳʤ˾˽äƲ̤γѤ롥12줬ͥǤꡤ6000쵬ϤĴǤ⤳268.7%ʡ#349. BNC Word Frequency List ˤ벻ʬĴ (2) ([2010-04-11-1]) ΥդȡˡоݤȤõϤˤꡤͥͭΨư뤳Ȥ狼뤬ŪˡѸäˤƤ1--3줬פǤ뤳Ȥϴְ㤤ʤ
ǤϡܸθäˤĤơ̤γϤɤƣۤ (80) Ǥϡˤܸ쥢ȼŵ٤θФ˴ŤʬۤĴ̤Ƥ롥ŵθФǤ뤫оθäϿεϤȻפ롥ʲΤ褦ʷ̤Ф
| 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | |
| 0.3 | 4.8 | 22.7 | 38.8 | 17.7 | 11.0 | 3.3 | 1.2 | 0.2 | 0.1 | 100 |
[2012-06-28-1], [2012-06-29-1]ϢҲ𤷤Ƥ MRC Psycholinguistic Database ˴Ťơ4ĤαѸפǡեλ˼ƤɽȤ˥դ̤Υѥ˴ŤĴԤʤäƤΤ⤢ΤǡӤͤͥǡϡHTMLȡ
(1) ʸˤ

(2) ǿˤ

ʻ͡
[2012-02-13-1]: #1022. ѸγƲǤ١
(3) ˤ

ʻ͡
[2010-04-09-1]: #347. ñʿѲϤɤΤ餤
[2010-04-10-1]: #348. BNC Word Frequency List ˤ벻ʬĴ
[2010-04-11-1]: #349. BNC Word Frequency List ˤ벻ʬĴ (2)
[2010-04-17-1]: #355. COLT Word Frequency List ˤ벻ʬĴ
(4) ʻˤ

ʻ͡
[2012-06-02-1]: #1132. ñʻ̤γ
[2011-02-23-1]: #667. COCA 50ʻ̤γϡ
[2011-02-22-1]: #666. COCA 5000ʻ̤γϡ
[2011-02-16-1]: #660. ѸΥեѸηƻΨ
¾ä٤䡤̤γˤĤƤϰʲε⻲ȡ
[2010-03-01-1]: #308. Ѹκѱñꥹȡ
[2011-08-20-1]: #845. Ѹθäεȳ
[2012-01-07-1]: #985. Ѹθäεȳ
εǡMRC Psycholinguistic Database 150837ޤˤѤ Amano θȤAmano Ǥϡ̾ư stress typicality ĴʪȤơƱǡ١˴Ťʻ̳ɽƤΤǡϤ⤷Ƥ
Amano (86) ϡǡ١10894Ĥ2ȴФʣʻεǽ碌ĸˤĤƤϡ줾ʻΤȤ1ĤȤƲäʤ¾ܤȼ p. 86 Ƥˡ̤Ȥ줿ʻ̤θĿȳϰʲ̤Ǥ롥
| POS | FREQ | % |
| noun | 7326 | 57.04% |
| verb | 2501 | 19.47% |
| adjective | 2420 | 18.84% |
| adverb | 291 | 2.27% |
| preposition | 68 | 0.53% |
| conjunction | 21 | 0.16% |
| pronoun | 15 | 0.12% |
| interjection | 37 | 0.29% |
| past participle | 57 | 0.44% |
| others | 108 | 0.84% |
ε[2012-05-05-1]˰³ grey eyes ꡥϡѸޥ grey eyes ˤĤƹͤζɽϸˤ³Ƥ롥BNCWeb ǡ"(grey|gray) {eye/N}" ȤƸȡ287㤬ҥåȤgrey eyes ̤ηƻԤƤߤȡclear, dark, deep, pale Ū¿beautiful bright 鷺ʤ餢ä
Τ褦㤫ȽǤȡgrey Τϵ̵ͭɽ魯̣ôƤʤ褦˻פ롥⤷ôƤȤСष pale ΡΤʤפȤ˰ѱѼdzǧ¤ꡤѸ grey ΰŪʸ촶ϡܸΤȤ褯ơnegative ϷǯµͫݵŷΥäơѸ grey eyes ϡnegative ʥ˥奢ä˴ްդʤɤߤȤʤС˿ȤƤΡֳפ뤤ϡĤߤΤ֤äפɽ魯Τȹͤ롥뤤ϡgrey eyes ϡ̣ޤäɽȤѤƤˤʤȤǽ⤢뤫⤷ʤ
ȡޤޤѸŵṲ̄ȤƤ grey eyes 狼ʤ⤷MED Silverstein Ҥ٤Ƥ̤ꡤѸ grey ɽ路ΤȤСѸεΤʤ grey 180٤ΰ̣ѲФȤˤʤ롥
gradation ΤǤꡤĤʤäƤϰϤ̣ꤷ뤳ȤϡʤʤѸΤߤʤ餺ܸˤƤ⡤̸Ǥ롥
ʤŵṲ̄ƤBrewer (258) ϡMatthew of Vandôme ˤ Helen of Troy ̤ʲ̤ꡤ1ĤηǤȤƤ롥
. . . her hair is golden, forehead white as paper, eyebrows black and thin. The space between the eyes (in contrast to the Greek ideal) is white and clear, a 'milky way'; the face is a shining star; the eyes are like stars. She has a little smile, a nose neither too big nor too small. Her face is rosy, her colouring white and red, like rose and snow. Teeth are like ivory, lips are small, slightly swelling, honeyed. Her mouth smells like a rose, her neck is smooth, shoulders radiant, well-spaced (dispatiati), breasts small, and figure incomparable.
ʽǤ礦ҲäƤߤ
Silverstein, Theodore, ed. Sir Gawain and the Green Knight. Chicago: U of Chicago P, 1983.
Brewer, D. S. "The Ideal of Feminine Beauty in Medieval Literature, Especially 'Harley Lyrics', Chaucer, and Some Elizabethans." The Modern Language Review 50 (1955): 257--69.
[2012-05-02-1], [2012-05-03-1]εǼ夲Ƥ Zipf's law ڡʤȤθˤ뤿ˡGeneral Service List (GSL) κ2000;ΥǡѤƷƤߤʥǡեˡ


ǽΥդٽ̤٤ݤ碌դǡٽ100̤ۤɤޤǤθоݤȤʲϤҤƤ椯ΤߤʤΤǾά٤ΥդޤǤʤѤοۤɤDZٸΤۤȤɤʤäƤޤͻҤ褯狼롥
ΥդϡZipf's law ˤˤʤȤٽ̤٤ѤļˤȤäΤǤ롥̿ޤǤϡפϾ岼礭ɤưꤷʤʸ1000줰餤ޤǤϡˤ䤫ϤΤΡ夯θΥճǤϤҤ³롥äơפǤΤܤ˸Ƥ1000줰餤ޤǤ
ˡ§ȸƤ֤ΤϤޤ˳Ƥȹͤ뤫Ū褯ФƤȤȤ館뤫ϡѻԤθҤȤĤǤ롥Zipf's law ˤפϡ֤褽פȲ᤹Τۤλ֤褽פɤ٤ǤΤƤʤޤZipf's law ĥƤΤȰۤʤꡤդ٤Ȥ륳ѥΥˤ¸褦
##1089,1090,1091 εǡؤ (information theory) θˤĤơä˸; (redundancy) ܤʤҲ𤷤ϡJakobson ˤ "Linguistics and Communication Theory" ꤹʸˤäơؤͿƤҥȤͤƤߤ
Jakobson ϡˡŪʲǤμŪħ (distinctive feature) ȡˤñ̤Ǥ "digit" 뤤 "bit" Ȥο˵Ťʹ¤˸ؤȾܤJakobson ξʬζФؤ줫ؤ٤뤳ȤϲξԤδ֤Ʊ뤷ƤϤʤȤϲȤȤƤ롥ä2δؿ˰ääΤǡҲ𤷤
(1) ϡäѤʪŪʾãθΨηϤλȤ (code) ˴ؿꡤȯԡԡʸ̮̣ϹθʤηϤ code ǤϤ뤬ϸưɬפȤ¦̤1Ĥˤcode Τߤܤ٤٤Ǥ롥code 1¦̤ˤʤȤϡ#1070. Jakobson ˤưԲķ6Ĥιǡ ([2012-04-01-1]) ǸȤǤ롥
There is a similar danger when interpreting human inter-communication in terms of physical information. Attempts to construct a model of language without any relation either to the speaker or to the hearer and thus to hypostasize a code detached from actual communication threaten to make a scholastic fiction from language. (250)
(2) ؤ (1) ռǡμˡѤƸηϤθΨ¬ȤȤΩηϤȤƤŪʸΨȡܤ٤θºݾθΨȤξƤʤФʤʤԤ type Ūlangue Ūʰ̣ǤθΨԤ token Ūparole Ūʰ̣ǤθΨȤФ狼䤹Jakobson ϡǤμŪħǤʤ֥ƥΩǵҤǤǽŪˤ "bit" ˤäƵҤǤȹͤƤꡤˤAȸBʸˡθΨʤɤӤǤȤƤ뤬ݲ줿ηϤȤƤ code θΨΤȤؤƤ롥ǡѤμºݤˤãθΨ¬ȤСܤνи٤̣νŤߤŤȤȤɬפǤ롥ȼºݤΥХפȤȤǤ롥
The amount of grammatical information which is potentially contained in the paradigms of a given language (statistics of the code) must be further confronted with a similar amount in the tokens, in the actual occurrences of the various grammatical forms within a corpus of messages. Any attempt to ignore this duality and to confine linguistic analysis and calculation only to the code or only to the corpus impoverishes the research. The crucial question of relationship between the patterning of the constituents of the verbal code and their relative frequency both in the code and in its use cannot be passed over. (251)
(2) ζθ츦˰ĤƲ᤹ȡ¤ؤȥѥؤϢȤȤ褦ʲˤĤʤäƤΤǤϤʤѥˤä줿ͤȤ˳Ƹܤ˽ŤߤŤԤʤΩνȤƵҤ줿ηϤΥѥȤƴޤƤ롥뤳ȤˤäơMartinet μĥηкθ ([2012-03-24-1], [2012-04-21-1]) ʤɤ⸡ڲǽȤʤΤǤϤʤ
Jakobson, Roman. "Linguistics and Communication Theory." Structure of Language and Its Mathematical Aspects. Providence: American Mathematical Society, 1961. 245--52.
ʸɤǤơthe choice is between rhyme or prose Ȥ˽Ф路between ˤ³ and ԤȤchoice θ촶˰ or ѤƤΤ餷˥缭ŵǤϡˡˤĤưʲΤ褦˿Ƥ롥
1(3) between 1980 to 1990 choose between war or peace Τ褦 and to or ѤΤ((ޤ))to from A to B䡥or choose, decide ʤɤưϢ줹Ȥ¿Ѥ롥 choice [decision] A or B ȹͤʢ2ˡ
2[̡ʬ] Ĥδ֤[]ĤΤɤ餫?choose ? peace and war ʿ¤褫Τ줫֡Ԣand or Ѥ뤳Ȥ롨 1 [ˡ](3)
OED Ǥϡ"between" 18 ̡ʬۤˡƤ뤬or ѤʸϵƤʤƱMED Ǥ bitwene 7 ˡб뤬Ϥ or ʸϤʤ"between A or B" 㤬ĸ줿ΤȤ䤤ˤϡܤ˥ѥĴ٤ɬפꤽ
ѸˤĤơBNCWeb ư "{choose/V} between_PRP + or_CJC" ȤƸʸʬȤۤ8ǤϤ뤬㤬줿 Written books and periodicals Ǥ롥Ū狼䤹4褦
. . . in 1627 Emperor Ferdinand ordered all his Bohemian subjects to choose between Catholicism or exile.
The main characters are all glorified psychopaths, with little to choose between hero or villain in terms of basic humanity.
. . . Mapleton, already out of breath, had to choose between talking or using his energy to keep up.
It is for you to choose between clinical or disciplinary action.
Ʊͤˡ̾ "{choice/N} between_PRP + or_CJC" θ̤⻲Ȥ줿
"between A or B" Ϥޤ˵ʹ¤餫ä˵ʸˡǹ⤵ƤǤʤԤ줬̡ꡤȽ̣ˤ or θ촶ˤ褯Ǥ뤷or λѤˤä¿Ǥ between θꤵΤ顤Τ褦ʸˡϤष侩٤ȹͤ롥
COCA ( Corpus of Contemporary American English ) Ĥ Mark Davies [2012-01-08-1]ε#986. COCA "WORD AND PHRASE . INFO"פǾҲ𤷤ǽ (Frequency List) ˲äʸꤲCOCA١dzƸ˴ؤŤ֤Ƥ륵ӥ WORD AND PHRASE . INFO, ANALYZE TEXT
Ŭʱʸꤲȡñ줬٥٥ˤäƿʬ줿֤֤롥500ޤǤĶٸġ3,000ޤǤιٸСʲ٤θϲǼۤacademic word ֻȤ֤롥ʸǤΤ줾γ⼨졤θåꥹȤФȤưפƸϥå֥ǡåΥץ뤬 KWIC DZڥɽ롥ޤڥˤ줬롥ʲϡε#1040. ̻ŪѲȶŪѰۡ ([2012-03-02-1]) ˰ѤʸꤲǤΥåȡ

ʸȤˤ collocation synonym Ĵ٤ʤȤ¿ΤǡȤǤϱѺʸؽ˰ϤȯʸϤ academic ٤ȽꤹΤˤȤ롥Academic Word List ˴ޤޤäδͭ٤ȤȤǤС[2010-12-30-1]ε#612. Academic Word Listפǵ The AWL Highlighter ġ
ε#1034. ѸˤɰդŪʡ ([2012-02-25-1]) (4) ǡѸǤϡ1;Τ¾;ΤˡŪɰաפ1;Τ֤뤳Ȥ˿줿ŪʸˡȤäƤ褤2;3;1;ΤȤ̤Ǥꡤ"you and I", "she and I", "you, he, and I" ʤɤȤʤ롥ΤȤƳؤȤϤޤºɤȸθǤꡤܸɤŨɰդθʤɤȴΤǤ롥Quirk et al. (Section 13.56, Note [a]) ˤϼΤ褦ˤ롥
When one of the conjoins is a personal pronoun, it is considered polite to follow the order of placing 2nd person pronouns first, and (more importantly) 1st person pronouns last: Jill and I (not I and Jill); you and Jill, (not Jill and you), you, Jill, or me (not me, you, or Jill), etc.
ƱݤεҤϡHuddleston and Pullum (1288), Biber et al. (338) ˤ⤢롥
ȤѸο;̾ν politeness Ȥ˻Ƥ褤Τɤ路ʤ뵭Ҥ˽Ф路ٹ (191) ˤȡʣǤ1;2;3;ΤνȤĤޤꡤ"we and you", "we and they", "we, you, and they" ʤɤȤʤ롥ǤϡɾˤϤʤʤ
ʣʷˤĤ BNCWeb Ĵ٤Ȥ⤽㤬ʤΤ褦ʤȤΤºݤΤȤ3Ĥο;Τʤɤϳ̵ä
we and you (0), you and we (0);
we and they (11), they and we (7)
you and they (11), they and you (6)
we, you, and they (0), we, they, and you (0)
you, we, and they (0), you, they and we (0)
they, we, and you (0), they, you, and we (0)
ʣˤĤƤϡʸ˭٤˵ΤȾȤٹˤʸäƤʤȤ餹ȡ餫εʸˡäƤΤʤΤ3緿ʸˡˤڤʤ
ñˤĤƤ⡤˼ϤޤǴǤꡤˤäƤϤδ㤫鳰⤢롥㤨СȤȤˤϡ1;Τ˽ФΤ褤Ȥ (ex. I and Bob were arrested for speeding.) ޤʬοʬΤۤ餫˾ξˤϡI and my children I and my dog ꤦ롥ѤϤȤƤ⡤ǽŪˤϥХ
Quirk, Randolph, Sidney Greenbaum, Geoffrey Leech, and Jan Svartvik. A Comprehensive Grammar of the English Language. London: Longman, 1985.
Huddleston, Rodney and Geoffrey K. Pullum. The Cambridge Grammar of the English Language. Cambridge: CUP, 2002.
Biber, Douglas, Stig Johansson, Geoffrey Leech, Susan Conrad, and Edward Finegan. Longman Grammar of Spoken and Written English. Harlow: Pearson Education, 1999.
ٹ ﵭرʸˡ3ǡʸƲ1926ǯ
cannot help doing ϡ?뤳ȤʤפȤ?ˤϤʤ?Τϻʤפ̣봷ɽǤ롥cannot but do ȤƤƱܿͤˤŪȤ䤹ɽɸΤ褦Ӥʸˤ than ΤʤǸƱʸˤդɬפǤ롥
ƤBNCWeb ˤ "(more (_AJ0 | _AV0)? | _AJC) * than * (can|could) (_XX0)? help" ǸȡϢ㤬8ҥåȤۤƱɽϺơ6
. . . the Commander struck out for the shore in a strong breaststroke that did not disturb the phosphorescence more than he could help . . . .
I'm not putting money in the pocket of the bloody Hamiltons more than I can help.
"Don't be more stupid than you can help, Greg!"
Resolutely, and determined to think no more than she could help about it . . . .
And I won't spend more than I can help.
"We'll do our best; we won't get in your way more than we can help."
ơιʸϡտޤƤ̣äƤˤ롥㤨СɤƤӡ3դϰޤˤʤͤФƤ̿ʸȯȡ3դޤǤϵ4դϰʡפȤݤȤʤʤǤä狼䤹뤿տϼȤˡʤȤ⡤줬ȯüԤΰտޤǤȹͤ롥Ū˹ͤȡyou can help ȹǤ뤫顤̤ϡȤޤˤ館뤮꤮̡4դؤϤ¿ϰʤȤȤ顤4դޤǤϵ5դϰʡפȤʤäƤޤĤޤꡤȯüԤΰտޤΰ̣ȤäƤޤޤŪˤΤǤС*Don't drink more pints of beer than you cannot help. ȤʤϤμιʸ BNCWeb Ǥʸڤʤ
Ǹо嵭Τ褦ˤʤ뤬Ԥΰտޤʸȯ뵡ϤۤȤɤʤ졤Ū˺뤳ȤϤʤޤ[2011-12-03-1]ε#950. Be it never so humble, there's no place like home. (3)פǸ褦ˡǤǤ̣ѤʤȤˤ狼ˤϿʤ褦칽¤Τ¸ߤ롥Ȥȡɸ칽¤ƤỤ̄Ū;ϤϤȤȤˤʤ롥
ʤߤˡɸʸϺǯλɸ1ĤǤ롥ˤĤƤϡġĤǤᤷƤ
Powered by WinChalow1.0rc4 based on chalow