Language

The company the word keeps. Every table here uses two measures side by side. Log-likelihood (written G²) says how confident we can be that a word turns up at a rate chance alone would not produce; log ratio says how large that difference is. Across 58,906,431 words almost anything reaches statistical significance, so confidence on its own is not a finding — the tables rank by confidence and report the size beside it.

The words that sit near a term

Which words appear near this term far more often than chance would put them there?

6,092 occurrences, 60,050 words in window
View the leading collocates as a table
WordNearLog ratio
crimes2,55817,826+6.52
humanity1,1109,474+7.74
war1,0935,317+4.93
rwanda5733,378+5.70
cleansing3593,196+8.04
ethnic3932,255+5.58
srebrenica2462,192+8.04
tutsi1631,898+10.66
rwandan1961,566+7.31
crime3271,380+4.44
denial1631,189+6.77
prevention2951,029+3.87
responsible248953+4.15
committed315926+3.44
punishment134882+6.23
adviser105617+5.69
twentieth86564+6.20
acts183517+3.35
atrocities118514+4.54
anniversary117510+4.54
perpetrators141479+3.80
victims177458+3.15
violations179376+2.75
ideology62372+5.78
perpetrated82358+4.55
protect146352+3.01
glorification42338+7.36
populations102324+3.62
prosecution74315+4.46
mass99301+3.52
Source

05_lexical.py → lexical/collocates.json

Download

The CSV holds every row behind this figure, not only what is drawn on screen, and each file names the script and word-list version that produced it.

The same table as a cloud

What shape does that neighbourhood have, and does it hold for a single speaker or a single decade?

nothing is drawn for a set of fewer than 20 speeches

Drawing…

40 words drawn, out of 100 held for the whole corpus: 0 fall below the minimum frequency and 60 fall beyond the number of words asked for. The table below holds those same 40 words, in the same order.

View the same words as a table
WordNearLog ratio
crimes3,29521,723+6.26
humanity1,51612,618+7.65
war1,5277,086+4.77
rwanda7524,174+5.46
cleansing4093,390+7.62
ethnic4582,360+5.16
srebrenica2732,248+7.58
tutsi1721,876+10.21
rwandan2191,604+6.84
prevention4361,470+3.78
crime3731,347+3.97
denial1761,153+6.23
committed395999+3.11
adviser165973+5.73
responsible292963+3.73
punishment137785+5.60
violations289620+2.79
glorification71590+7.63
perpetrators184561+3.53
twentieth93541+5.67
atrocities136512+4.09
acts211479+2.89
victims220477+2.81
protect210473+2.88
anniversary128469+4.01
criminals89427+4.89
prosecution102408+4.27
serious207367+2.46
perpetrated92342+4.06
ideology65338+5.20
tutsis32316+9.07
impunity148315+2.77
justice230308+2.07
responsibility212303+2.15
prosecute67292+4.55
mass114283+3.07
convicted51274+5.33
populations109269+3.06
dieng32261+7.51
commemoration40256+6.12
Source

05_lexical.py → lexical/collocates.json

Download

The CSV holds every row behind this figure, not only what is drawn on screen, and each file names the script and word-list version that produced it.

The same word in two mouths

Do different speakers, groups or decades use it to do different work?

Align
nothing is drawn for a set of fewer than 20 speeches

Rwanda 187 speeches · 960 occurrences

  1. rwanda +6.64
  2. tutsi +11.79
  3. denial +6.85
  4. convicts +10.39
  5. ideology +7.06
  6. fdlr +7.39
  7. fugitives +6.38
  8. rwandan +6.61
  9. suspects +6.68
  10. masterminds +10.41
  11. tutsis +10.26
  12. twentieth +6.52
  13. perpetrators +4.26
  14. crimes +3.23
  15. survivors +5.31
  16. post-genocide +13.45
  17. crime +3.73
  18. committed +3.13

United States 163 speeches · 322 occurrences

  1. crimes +5.79
  2. humanity +6.93
  3. srebrenica +8.61
  4. rwandan +8.01
  5. war +4.12
  6. rwanda +5.07
  7. atrocities +5.40
  8. kabuga +8.20
  9. responsible +4.29
  10. accused +5.94
  11. financier +12.69
  12. committed +3.31
  13. felicien +9.06
  14. ethnic +4.41
  15. darfur +3.71
  16. denial +5.87
  17. alleged +5.34
  18. violations +3.10
View both profiles as a table
ProfileWordNearLog ratio
Rwandarwanda2711,951+6.64
Rwandatutsi1221,666+11.79
Rwandadenial45337+6.85
Rwandaconvicts24290+10.39
Rwandaideology37288+7.06
Rwandafdlr31255+7.39
Rwandafugitives36246+6.38
Rwandarwandan33236+6.61
Rwandasuspects30218+6.68
Rwandamasterminds17206+10.41
Rwandatutsis16191+10.26
Rwandatwentieth27190+6.52
Rwandaperpetrators47188+4.26
Rwandacrimes69186+3.23
Rwandasurvivors33178+5.31
Rwandapost-genocide11165+13.45
Rwandacrime49162+3.73
Rwandacommitted62159+3.13
United Statescrimes139843+5.79
United Stateshumanity63480+6.93
United Statessrebrenica37367+8.61
United Statesrwandan30273+8.01
United Stateswar53203+4.12
United Statesrwanda32163+5.07
United Statesatrocities18100+5.40
United Stateskabuga1094+8.20
United Statesresponsible2393+4.29
United Statesaccused1381+5.94
United Statesfinancier576+12.69
United Statescommitted2467+3.31
United Statesfelicien663+9.06
United Statesethnic1563+4.41
United Statesdarfur1653+3.71
United Statesdenial849+5.87
United Statesalleged949+5.34
United Statesviolations1948+3.10
Source

05_lexical.py → lexical/collocates_sliced.json

Download

The CSV holds every row behind this figure, not only what is drawn on screen, and each file names the script and word-list version that produced it.

Compared with a like-for-like speech

Setting aside what the debate was about, what marks out a speech that says genocide?

3,104 of 3,273 speeches found a partner (94.84%)
#WordIn these speechesLog ratio
1genocide5,1875,306+12.76
2crimes8,1622,159+1.56
3rwanda4,6241,051+1.41
4humanity2,093936+2.33
5justice5,272458+0.78
6court3,553445+0.96
7war4,901357+0.70
8atrocities1,291332+1.53
9genocidal295302+8.62
10rwandan784299+2.04
11criminal3,727283+0.72
12srebrenica549240+2.29
13victims2,529240+0.82
14cleansing722235+1.82
15impunity2,189235+0.88
16accountability2,044208+0.85
17mass924201+1.37
18ethnic1,488200+1.01
19committed3,084180+0.62
20atrocity388170+2.29
21icc1,136165+1.06
22ictr1,333161+0.94
23responsible1,478155+0.87
24tutsi184146+4.36
25violations2,446142+0.61
26rights5,410138+0.39
27sexual2,523137+0.59
28arrest870133+1.09
29prevention1,866133+0.69
30human5,433129+0.38
Source

05_lexical.py → lexical/keyness.json

Download

The CSV holds every row behind this figure, not only what is drawn on screen, and each file names the script and word-list version that produced it.

Which terms travel together

Does this vocabulary hold together in groups, or is it just a list?

a line is drawn where at least 20 speeches use both terms
Source

05_lexical.py → lexical/network.json

Download

The CSV holds every row behind this figure, not only what is drawn on screen, and each file names the script and word-list version that produced it.

Every word above is a way in: the concordance holds all 3,273 speeches these tables were built from.

Top