hellogѸ˥֥     ChangeLog ǿ    

lltest - hellogѸ˥֥

ǽ: 2026-07-15 01:27

2012-10-31 Wed

#1283. η׻ˡ [corpus][statistics][bnc][collocation][lltest]

[2010-03-04-1]ε#311. girl Ȥ褯 collocate ƻϲפǡȸζ (collocation) ¬׻ˡ (association measure) ˤϤĤμब뤳Ȥ򸫤ѥؤǤϡLog-Likelihood Test Ȥˤ׻ˡŪ褯ȤƤ뤬줾η׻ˡˤħΤǡʤ٤ʣˡΤ褤[2010-03-04-1]ƤȽʣʬ⤢뤬BNCweb ǼƤ7η׻ˡγơˤĤ Hoffmann et al. (149--58) 򻲾Ȥʤ顤ħѤΥҥȤ򼨤
Ƽη׻ˡϡ(a) (frequency of co-occurrence)(b) ͭ (significance of co-occurrence)(c) եȡ (effect-size) 1ġ뤤ʣȤ߹碌˴ŤƤ롥(b) ϡŪͭդǤȤγο٤ɽ魯ɸǤꡤζɽ魯ΤǤϤʤȤդɬפ롥(c) ϡѻ٤ȴ٤Ȥ׻δܤȤɸǤ롥

(1) Rank by frequency
ѻ붦٤ΤΤѤ롤ǤñľŪʼ١¾η׻ˡΤ褦ʣ׽ϤۤɤƤ餺ɸȤƤϺǤƤǽɵʤɤ̤뤳Ȥ¿̾ζʬϤˤѤʤ

(2) Log-likelihood
ͭѤ롥BNCweb ΥǥեȤη׻ˡǡѥǹѤƤ롥ǽɵʤɤζˤƹ٤θȤζ䡤դ˶ˤ٤θ1, 2ʤɡˤȤζϤ롥٤ι⤤Ȥ߹碌˹ͿȤħꡤˤդפ롥

(3) Mutual information (MI)
եȡѤ롥ˤ褯ѤƤ׻ˡѤäƤ¿դפ롥ǽɵʤɤȤΤդ줿ŪӽƤϤ褤ȿ̡٤ζɽؤФ꤬㤷Фαƶ򸺤뤿ˡBNCweb Ǥ "Freq(node, collocate) at least" 10ʾꤹ뤳Ȥ侩롥ˤꡤ"conspicuous and intuitively appealing collocations involving words of intermediate frequency" (Hoffmann et al. 154) ⤭ĦȤʤ롥

(4) T-score
٤ȶͭθ׻ˡ٤1ʲ٤εʶɽˤĤƤ Rank by frequency Ȼ褦ʿ񤤤򤷡٤ι⤤ɽˤĤƤ϶ͭȿǤ񤤤򤹤롥ޤѻ٤٤ɬ⤯ʤ롥Log-likelihood ̤Ȥʤ뤳Ȥ¿٤ؤΥХϰضʤ롥ΡɤΤΤ1000礭ˡ̤ȯ뤳Ȥ롥

(5) Z-score
ͭȥեȡθ׻ˡ٤ζɽˤϥեȡŻ뤹뤬٤ζɽˤϤޤǥեȡ˴꤫ʤLog-likelihood MI ξħ褦ʡХ󥹤μ줿ɸǤ롥MI Ʊͤˡ٤ζɽؤΥХߤΤǡ"Freq(node, collocate) at least" 5٤ꤹΤ褤Ȥ롥

(6) MI3
٤ȥեȡθ׻ˡMI ΤɽؤнŤ٤Ƥ롥ٶɽˤϥեȡٶɽˤ϶٤Ū褯ȿǤ롥POS ˤȤȤѤȸŪʣ줫ʤѸʤɤμФ˰Ϥȯ롥ΤȤƤϹٶɽؤΥХŪʶʬϤˤϸʤ

(7) Dice coefficient
MI3 Ʊͤˡ٤ȥեȡθ׻ˡMI3Ȱۤʤꡤٶɽˤ϶٤ٶɽˤϥեȡ褯ȿǤ졤ξԤڤؤޤʤΤħŪǤ롥ڤؤϡΡɤΤΤ٤ɽ٤10ܤۤɤǵȤ롥иŪˡZ-score Ȼ褦ʷ̤뤬Z-score ۤ٤˴ŤХʤ

ʾΤ褦¿ढäܰܤꤹ뤬Hoffmann et al. θˤСñη׻ˡȤƤ Log-likelihood MI ǡη׻ˡȤƤ Z-score Dice ȤΤȤǤ롥
͡ʷ׻ˡˤĤƤϡAssociation measures 򻲾ȡ

Hoffmann, Sebastian, Stefan Evert, Nicholas Smith, David Lee, and Ylva Berglund Prytz. Corpus Linguistics with BNCweb : A Practical Guide. Frankfurt am Main: Peter Lang, 2008.

Referrer (Inside): [2019-07-10-1]

[ | ѥڡ ]

2012-10-26 Fri

#1278. BNC 濴Ȥ륳ѥϢΥ󥯽 [corpus][bnc][link][web_service][lltest]

ѥؤߤޤʤʬʬˡϢϥ־뤳Ȥ¿ʤ褦ˤפ뤬դ˾¿ơȽǤ˺롥ƼʬΤǤʥ󥯽ޤȤƤȻפΤسΥԡɤˤĤƹԤʤ䤬Ǥ褯Ѥ BNC ˴ϢΤ濴ˡŪǤϤ뤬󥯤ĥ롥󥯽ޤȤϫ򼹤ϡŤ뼰ˤɤ뤫ɸΤۤΨŪȤˤʤĤĤ롦

1. BNC 󥿡ե

BNCweb ̵Ͽ
BYU-BNC ̵Ͽ
BNC ( The British National Corpus )

2. BNC Υե󥹡

Quick Reference for Simple Query Syntax (PDF)
Reference Guide for the British National Corpus (XML Edition)
Reference Guide ܼ
* 6.5 Guidelines to the Wordclass Tagging
* The BNC Basic (C5) Tagset
* 9.8 Simplified Wordclass Tags
* 9.7 Contracted forms and multiwords
* 1 Design of the Corpus
* 9.6 Text and genre classification code

3. ѥϢ祵

David Lee ˤ Bookmarks for Corpus-based Linguists
* Corpora, Collections, Data Archives
* Software, Tools, Frequency Lists, etc.
* References, Papers, Journals
* Conferences & Project

4. hellog ε

#568. ѥȱѸ쥳ѥ: [2010-11-16-1]
#506. CoRD --- Ѹ˥ѥξ󥻥󥿡: [2010-09-15-1]
#308. Ѹκѱñꥹȡ: [2010-03-01-1]
ѥϢ: corpus
BNC Ϣ: bnc
COCA Ϣ: coca

5. ׻ġ

Corpus Frequency Wizard
Paul Rayson's Log-likelihood Calculator
VassarStats
hellog Ρ#711. Log-Likelihood Tester CGI, Ver. 2: [2011-04-08-1]

Hoffmann, Sebastian, Stefan Evert, Nicholas Smith, David Lee, and Ylva Berglund Prytz. Corpus Linguistics with BNCweb : A Practical Guide. Frankfurt am Main: Peter Lang, 2008.

Referrer (Inside): [2015-04-22-1]

[ | ѥڡ ]

2011-10-28 Fri

#914. BNC ˤä庹Ĵ [bnc][corpus][statistics][lltest][interjection]

ε#913. BNC ˤä˽Ĵ ([2011-10-27-1]) Ǽꤢ Rayson et al. ǤϡüԤ̤ǤʤǯˤäѰۤĴƤ롥ǯ𺹤ȤäƤ⡤35̤ʾ夫Ǿ岼ʬ绨Ĥʬ̤ϤĤζ̣ͿƤ롥ʲϡχ2 ξ19̤ޤǤΰǤ (142--43)

RankUnder 35Over 35
Wordχ2Wordχ2
1mum1409.3yes2365.0
2fucking1184.6well1059.8
3my762.4mm895.2
4mummy755.2er773.8
5like745.2they682.2
6na as in wanna and gonna712.8said538.3
7goes606.6says443.1
8shit410.1were385.8
9dad403.7the352.2
10daddy380.1of314.6
11me371.9and224.7
12what357.3to211.2
13fuck330.1mean155.0
14wan as in wanna320.6he144.0
15really277.0but139.0
16okay257.0perhaps136.0
17cos254.4that131.3
18just251.8see122.1
19why240.0had118.3


ͽ̤ۤꡤ㤤ħŪʥɤϤ¿ޤǤ롥ɽθޤƤyeah, okay, ah, ow, hi, hey, ha, no, ooh, wow, hello ʤɤδ졤fucking, shit, fuck, crap, arse, bollocks ʤɤΥ֡줬Ωġ㤤ΥɤȤơ츫ͽۤ󤬤롥㤨Сplease, sorry, pardon, excuse ʤɤǫ줬㤤ħŪȤ
ۤˤϡ㤤ħŪʷƻ줬Ĥ (ex. weird, massive, horrible, sick, funny, disgusting, brilliant, really, alright, basically) ɾɽ魯ƻ졦줬¿ήԤȤߤʤȤǤ췲ǯ𺹤 "apparent time" κȹͤСˤ "real time" Ѳ뤳ȤˤʤΤǡθ췲̻Ū٤äõΤ⤪⤷

Rayson, Paul, Geoffrey Leech, and Mary Hodges. "Social Differentiation in the Use of English Vocabulary: Some Analyses of the Conversational Component of the British National Corpus." International Journal of Corpus Linguistics 2 (1997): 133--52.

Referrer (Inside): [2013-04-14-1] [2011-11-02-1]

[ | ѥڡ ]

2011-10-27 Thu

#913. BNC ˤä˽Ĵ [bnc][corpus][statistics][lltest][interjection][gender_difference]

ɸ򰷤ä Rayson et al. ʸɤBNC ǡ͸Ūʴʬव줿äդϿ֥ѥ4,552,555ˤоݤȤơä˽ǯ𺹡ҲŪϰ̤ˤ뺹餫ˤ褦ȤǤ롥װΤʤǡŪѰۤŪ˺Ǥ⶯줿Τˤ뺹äȤȤʤΤǡܵǤϤη̤Ҳ𤷤
ޤʲ˵󤲤ͤβˤμɬפʤΤǡ˿ƤBNC ˼Ͽ줿äդϻִԤ2֤μʲä Walkman ˿᤭Ǥäǡ񤭵ΤǤꡤλִԤ73̾75̾Ǥ롥äо줹ִ԰ʳüԤˤĤƤ⡤Τۤ¿äơ֥ѥؤλΨǤСΤȤƽ⤯ʤ뤳ȤԻ׵ĤǤϤʤ
ƧޤǤ⡤ΤȤƽΤۤ褯äȤȤ򼨺ͤФѤ줿 word token ǤС1.00ȤȽ1.51äͭΨǤϡ1.00ȤȽ1.33ä˽βäǤΤۤ⤤ͭΨ򼨤ȤԸ椬뤬BNC Υ֥ѥǤϽƱΤβä¿äȤȤ嵭η̤طʤˤΤ⤷ʤˤ衤̣ͤǤ뤳Ȥϴְ㤤ʤ
ˡ٤äˤ˽򸫤Ƥߤ褦˽ٹ礤ι⤤ɤȴФˡϡȤƤ[2010-03-10-1], [2010-09-27-1], [2011-09-24-1]εǾҲ𤷤ΤƱˡǤ롥ѥȽѥ̤줾줫줿ɽͤ碌Ū˽ (χ2) ι⤤¤ؤФ褤ʲϡ25̤ޤǤΰǤ (136--37)

RankCharacteristically maleCharacteristically female
Wordχ2Wordχ2
1fucking1233.1she3109.7
2er945.4her965.4
3the698.0said872.0
4year310.3n't443.9
5aye291.8I357.9
6right276.0and245.3
7hundred251.1to198.6
8fuck239.0cos194.6
9is233.3oh170.2
10of203.6Christmas163.9
11two170.3thought159.7
12three168.2lovely140.3
13a151.6nice134.4
14four145.5mm133.8
15ah143.6had125.9
16no140.8did109.6
17number133.9going109.0
18quid124.2because105.0
19one123.6him99.2
20mate120.8really97.6
21which120.5school96.3
22okay119.9he90.4
23that114.2think88.8
24guy108.6home84.0
25da105.3me83.5


ɬ⤳25̤ޤǤɽǤɤ߼ʤRayson et al. (138--40) ˤаʲܤͤȤ

"four-letter words"졤δħŪǤ (ex. shit, hell, crap; hundred, one, three, two, four; er, yeah, aye, okay, ah, eh, hmm)
;̾졤1;̾졤δϽħŪǤ (ex. she, her, hers; I, me, my, mine; yes, mm, really) ̾λѤˤä˽Ϥʤ
the of λѤ¿˰̾Ѥ̾λѤ¿Ȥ̤λ¤ȴϢ뤫
ͭ̾졤̾졤ưϽ¿λ "report" ηФδط "rapport" ηθ줫
ͭ̾ΤʤǤ⡤̾ϽλѤ¿̾λѤ¿

¾Υѥˤ븡ڤɬפη̤Ȳ˶̣ߤ뤳ȤϳΤǤ롥
ɤ׽ȴϢơѥؤǥ踡ѤȤƹѤ褦ˤʤäƤ Log-Likelihood ˤĤƤϡ Log-Likelihood Tester, Ver. 1 Log-Likelihood Tester, Ver. 2 򻲾ȡ

Rayson, Paul, Geoffrey Leech, and Mary Hodges. "Social Differentiation in the Use of English Vocabulary: Some Analyses of the Conversational Component of the British National Corpus." International Journal of Corpus Linguistics 2 (1997): 133--52.

[ | ѥڡ ]

2011-04-08 Fri

#711. Log-Likelihood Tester CGI, Ver. 2 [corpus][bnc][statistics][web_service][cgi][lltest]

ʲˡѤ Log-Likelihood Tester, Ver. 2 ʸ褦ˡϥǡΥեޥåȤ䡤⡼ɤŬڤ򤵤ƤʤˤϥСǥ顼ǽΤա

each-line mode lump mode


[2011-03-25-1]εǡѥǤ褯Ѥпٸ ( Log-Likelihood Test ) η׻ Log-Likelihood Tester, Ver. 1 Ver. 1 ϡѥ̣ʤ2ĤΥѥǤΥɡʷˤνи٤١ѥ֤κͭդǤ뤫ɤꤹΤä
Log-Likelihood Test ϾҤŪѤ뤳Ȥ¿ȻפVer. 1 ǤϤƵǽòΤŪʣԡʣʬɽͿǡбпٸԤʤ⤢롥㤨Сε[2011-04-07-1]ǡѸˤ though although νиˤĤ BNC ˴ŤĴҲ𤷤Text Domain ȤΨϡξδ֤ŪˤɤٰפƤ롤뤤ϰפƤʤȤߤʤȤǤΤΥդ顤although ϳؽѻʸ¿though Ϻʸ¿ȤľŪʡְפŪˤϤɤΤ褦ɽΤ
Τ褦ʾˤϡΤ褦ɽͤ100νи٤ɸಽѤߡˤ򥳥ԡϥܥåŽդ롥"lump mode" ˥åؤ"Go!" 롥ʥǥեȤ "each-line mode" ǡ Ver. 1 ƱΥ⡼ɡ

    thoughalthough
Natural and pure sciences56.380.13
Applied science37.3668.31
World affairs45.8168.2
Social science48.9863.38
Commerce and finance46.1857.21
Arts74.0752.93
Leisure45.8549.46
Belief and thought70.7846.75
Imaginative prose80.226.37


̤ϡ1ԤɽȤƽϤ롥though although ɽ魯2οͤ¤ӤŪˤɤΤ餤Ƥ뤫׻Ƥ롥ȤƤϡξ Text Domain Ȥ٤¤Ӥκ p < 0.0001 Ȥ˹⤤٥ͭդǤꡤξνи Text Domain ˤäƤۤܳμ¤˰ۤʤȤ롥
ϥܥåǡν񼰤ϡֶڤʬɽɽƬɽ¦ϤάġץΤ褦ɽƬɽ¦ξޤˤϡΥ϶ˤƤɬפꡥ
"each-line mode" εǽ Ver. 1 ȸߴʤΤǡϷ⤽򻲾ȡ Ver. 2 "each-line mode" ǤϡϷ̤򥷥ץˤƤʵդˡܤ׻ͤˤ Ver. 1 Τۤͭѡˡ
Log-Likelihood Test γפˤĤƤϡ[2011-03-24-1]ε򻲾ȡ

Referrer (Inside): [2012-10-26-1]

[ | ѥڡ ]

2011-04-07 Thu

#710. though although θˡκ (2) [bnc][corpus][lltest][conjunction][statistics]

ε[2011-04-06-1]ǡthough although θˡκ˿줿Ʊǡ
4000Ķʤ The Longman Spoken and Written English Corpus (the LSWE Corpus) ȤѸʸˡBiber et al. (845--46) ǤϼΤ褦ˤ롥

Both of these subordinators [though and although] occur in all four registers [conversation, fiction, news, and academic prose], although the registers show different preferences of use. Conversation and fiction show a slightly greater use of though (concessive clauses are, however, uncommon in conversation generally). News shows no particular preference. In academic prose, although is about three times as frequent as though. Although seems to have a slightly more formal tone to it, fitting the style of academic prose . . . . The greater use of although by writers of academic prose may also result from an attempt to distinguish this subordinator from the common use of though as a linking adverbial in conversation . . . .


ޤƱ p. 842 ɽϡŪ though fiction ¿although academic prose ¿Ȥǧ롥ˤ뺹ƤȤη̤
Τ褦Ըơ BNC ( The British National Corpus ) ˤꤳΤƤߤ롥BNCweb ǡ{although/CONJ}, {though/CONJ} 򤽤줾측Written/Spoken, Text Domain, Sex of Author/Speaker, Perceived Level of Difficulty ʤ͡ʥѥ᡼ǽиʬۤʬϤΩä̤ʲ˼ʿͥǡϤΥڡHTML򻲾ȡˡ
ޤWritten/Spoken κˤĤƤϡͽۤȤꡤξȤ Written ؤФ꤬㤷ʺ۷ though 0.66344 although 0.49770 ǡ餫˽񤭸դФˡLog-Likelihood Test Ǥϡp < 0.0001 Υ٥ǽ񤭸դäդͭպΤ˼줿
񤭼ꡤäˤ뺹ⶽ̣񤭸դäդξǡalthough ͭպäλѤФäƤ롥though ˤĤƤϡ although ۤɸǤϤʤʤ񤭸դǤ p < 0.05 ͭպˡ
ˡText Domain ̤٤ߤ롥9 Text Domain ̤ ( Natural and pure sciences, Applied science, World affairs, Social science, Commerce and finance, Arts, Leisure, Belief and thought, Imaginative prose ) 100νиɸಽͤǡξ Text Domain ٤򥰥ղΤʲοޤ



Text Domain ˤäξνи٤оŪʷ뤳Ȥ狼롥Ū sciences ( = academic prose ) although ΩImag(inative) Prose ( = fiction ) though ¿Log-Likelihood Test ǤϡText Domain ˤиκ p < 0.0001 ͭդǤ롥
ľŪˤԸη̤ͽۤȤǤϤ뤬although ν񤭼ˤؽѻʸǸѤȤ޼줿

Biber, Douglas, Stig Johansson, Geoffrey Leech, Susan Conrad, and Edward Finegan. Longman Grammar of Spoken and Written English. Harlow: Pearson Education, 1999.

Referrer (Inside): [2011-04-10-1] [2011-04-08-1]

[ | ѥڡ ]

2011-03-25 Fri

#697. Log-Likelihood Tester CGI [corpus][bnc][statistics][web_service][cgi][lltest][sociolinguistics]

ε[2011-03-24-1] Log-Likelihood Test ˤ׻ˤ Rayson Log-likelihood calculator ѤФ褤ȽҤ٤ºݤθκݤ˺Ȥ⤦ưȻפäΤ CGI 򼫺Ƥߤ٤ϤȻפȤꤢ



ΥƥȥܥåϤ٤ǡϡֶڤɽη1ܡʾάġˤϥѥ̾2ܰʹߤϥɤȴѻٿʥҥåȿˡǽԤϳƥѥΥʸˡ"#" ǻϤޤԤϥȹԤȤ̵뤵롥1ܤΥϾάġ
ʲΥƥȤϥץ롥[2010-09-11-1]εǼ夲ƥӹѤƻӵȺǾޤ˥ȥå20٤BNCweb äե֥ѥüԤ̤ɽǤ롥ΤޤޥԡϥܥåŽդȡϷ̤ǧǤ롥

    BNC_Male_SpeakersBNC_Female_Speakers
new14991
good408310
free17375
fresh84118
delicious1234
full210107
sure532328
clean197223
wonderful270258
special17782
crisp1016
fine347215
big470415
great20396
real16380
easy326157
bright113110
extra347203
safe18292
rich12045
#--------
corpus_size49499383290569


˽֤ͭպä礭ΤϡбԤ֤ɤĤ֤줿 fresh, delicious, clean, wonderful, big ǡٿ˴ŤƷ׻줿 Diff_Co ( "Difference Coefficient" ֺ۷ ) ޥʥǤ뤳Ȥ顤ħŪʷƻȤȤˤʤ롥big ϰճʵ⤷̤Ǥ롥Фäͭպ򼨤ΤϲǼ easy rich Ǥ롥η̤Ϥɤ߹ळȤǤܺ٤Ĵ٤뤳ȤǤ롥ηƻȤϡüԤǤϤʤʹ̡ǯ𡤼Ҳ񳬵ʤɤ򼴤ĴƤ⤪⤷ȱѤǤ롥

Referrer (Inside): [2011-04-08-1]

[ | ѥڡ ]

2011-03-24 Thu

#696. Log-Likelihood Test [corpus][bnc][statistics][lltest]

[2010-03-04-1]εǿ줿ѥؤǤϳƼ׼ˡѤ롥ĤˡΤʤǤ⡤ɽΥѥ֤٤Ӥꡤcollocation ٹ礤¬Τ˹ѤƤΤ Log-Likelihood Test ( LL Test, G Test, G2 Test ʤɤȤ˸ƤФ븡Ǥ롥ѥθ줿ʤΤǥΰۤʤ륳ѥ֤ǤӤǽǤꡤƱŪǰˤ褯ѤƤ2踡 ( Chi-Squared Test ) ⤤ĤǤ줿ˡɾƤꡤǶΥѥǤϹѤƤ롥㤨С2踡ϴ٤5꾯ʤȤٸ򰷤Ȥѥ礭ΤȾΤӤȤ˿㤯ʤ뤬Log-Likelihood Test Ϥαƶˤ [ Rayson and Garside 2 ]
Log-Likelihood Test δŪʹͤϡѥȤˤɽδԤи١ʴ١ˤФͤȼºݤ˽и١ʴѻ١ˤκñʸȹͤۤɤ˶Ƥ뤫ɤȽꤹȤΤǤ롥ȤơΤ褦ʥǥBNC ( The British National Corpus ) äե֥ѥȽ񤭸ե֥ѥ̤ξ֥ѥ֤ f*ck Ȥ four-letter word ٤Ӥ롥BNCweb ꤳΥɤ򸡺ȡΤ褦ʷ̤줿

CategoryNo. of wordsNo. of hitsDispersion (over files)Frequency per million words
Spoken10,409,85857963/90855.62
Written87,903,571743172/3,1408.45
total98,313,4291,322235/4,04813.45


׽ۤɤޤǤʤDZ "Frequency per million words" 򸫤Сf*ck Ūäդ¿Ѥ뤳Ȥʬ뤬ϤŪ΢դ롥ޤ̵Ȥơäե֥ѥȽ񤭸ե֥ѥδ֤Ǥ f*ck ٺϸϰǤꡤθ˴ؤξԤ˰̣Τ뺹Ϥʤפꤹ롥Ωϡäե֥ѥȽ񤭸ե֥ѥδ֤Ǥ f*ck ٺϸϰǤʤθ˴ؤξԤκϰ̣פȤʤ롥̵⤬ٻ뤫ɤΤŪǤ롥

 Corpus 1Corpus 2Total
Frequency of wordaba+b
Frequency of other wordsc-ad-bc+d-a-b
Totalcdc+d


Log-Likelihood Test Ѥ Log-Likelihood ratio пפϡɽΤdzƥ֥ѥ ( c, d ) ȡƥ֥ѥǤ f*ck ٿ ( a, b ) ʬɽˤޤȤ᤿ǡ줾δ E1 E2 򲼤 (1) μǵᡤͤ (2) μƵ롥

(1) E1 = c*(a+b)/(c+d); E2 = d*(a+b)/(c+d)
(2) LL = 2*((a*log(a/E1))+(b*log(b/E2)))

f*ck οͤǷ׻ȡʲΤ褦ˤʤ롥

E1 = 10409858*(579+743)/(10409858+87903571) = 139.979170861796
E2 = 87903571*(579+743)/(10409858+87903571) = 1182.0208291382
LL = 2*((579*log(579/139.979170861796))+(743*log(743/1182.0208291382))) = 954.2115

Log-likelihood ratio Ȥ 954.2115 ȤͤФ롥ˤͤŬڤͭտ̾ 5%, 1%, 0.1%ˤб륫ͤӤ롥2 * 2 ʬɽФ׻Ǥϼͳ1ΥͤѤ뤳ȤˤʤäƤꡤͤͭտ 5%, 1%, 0.1% νˤ줾 3.84, 6.63, 10.83 Ǥ롥954.2115 Log-Likelihood ratio ͭտ 0.1% б 10.83 ⤺äȹ⤤Τǡ0.1% ͭտǵ̵ϴѤ롥СŪˤϵ̵⤬ǤΨ 0.1% ˤޤȹͤƤ褤ȤȤǤ롥Τ褦ˤΩäե֥ѥȽ񤭸ե֥ѥδ֤Ǥ f*ck ٺϸϰǤʤθ˴ؤξԤκϰ̣פ򤵤뤳Ȥˤʤ롥
Log-Likelihood Test ϰʾΤ褦˿ʤ뤬θԤʤˤäƤΤäƤɬפ롥̤ˤϡ׻٤ 5 򲼲륻뤬1ĤǤ⤢ˤϡ٤Ȥ롥 the Cochran rule ȸƤФƤ뤬꤭٤ʥ롼󵯤 Rayson, Berridge, and Francis (8) ˤС٤٤ͤͭտ 5% 13 1% 11 0.1% 8 Ȥͭտ 0.01% ꤹд 1 ˤѤ٤ΤǡRayson et al. ϥѥؤǴŪѤƤ3Ĥο˲äơ0.01% οб륫ͤ 15.13 ˤޤǤθ侩Ƥ롥
פˤϾܤʤɽ 2ʥ֡˥ѥ֤ǤӤȤǴñѤ뤳ȤǤ븡ȤơLog-Likelihood Test αϰϤϹ׻Τ Rayson Log-likelihood calculator ʤɤǤФ褤ܵϤΥڡεҤȥʸ򻲹ͤˤˡ
BNC Ѥ f*ck ϢʬۤθϡMcEnery et al. (264--86) Υǥ˾ܤ
ϢơϹԤʤʤäĤܥ֥ǰä gorgeous Ĵ ([2010-08-16-1], [2010-08-17-1],[2010-12-25-1]) ʤɤ⻲ȡ

Rayson, P., D. Berridge , and B. Francis. "Extending the Cochran Rule for the Comparison of Word Frequencies between Corpora." Le poids des mots: Proceedings of the 7th International Conference on Statistical Analysis of Textual Data (JADT 2004), Louvain-la-Neuve, Belgium, March 10-12, 2004. Ed. Purnelle G., Fairon C., and Dister A. Louvain: Presses universitaires de Louvain, 2004. 926--36. Available online at http://www.comp.lancs.ac.uk/computing/users/paul/publications/rbf04_jadt.pdf .
Rayson, P. and R. Garside. "Comparing Corpora Using Frequency Profiling". Proceedings of the Workshop on Comparing Corpora, Held in Conjunction with the 38th Annual Meeting of the Association for Computational Linguistics (ACL 2000), 1-8 October 2000, Hong Kong. 2000. 1--6. Available online at http://www.comp.lancs.ac.uk/computing/users/paul/phd/phd2003.pdf .
McEnery, Tony, Richard Xiao, and Yukio Tono. Corpus-Based Language Studies: An Advanced Resource Book. London: Routledge, 2006.

[ | ѥڡ ]

Powered by WinChalow1.0rc4 based on chalow