{"id":311,"date":"2013-02-06T22:13:28","date_gmt":"2020-02-18T06:53:26","guid":{"rendered":"https:\/\/redfrontdoor.org\/blg2\/?p=311"},"modified":"2020-02-19T15:03:31","modified_gmt":"2020-02-19T15:03:31","slug":"post-305","status":"publish","type":"post","link":"https:\/\/redfrontdoor.org\/blog\/?p=311","title":{"rendered":"Most cliched adjectives and nouns"},"content":{"rendered":"<div style=\"float: right; margin-left: 2em;\"><img decoding=\"async\" src=\"\/blog\/wp-content\/uploads\/2013\/02\/letter-tiles-cliche.jpg\" \/><br \/>\n<span style=\"font-size: 80%;\"><a href=\"http:\/\/www.flickr.com\/photos\/noodle93\/4652373324\/in\/photostream\/\">Photo<\/a> \u00a9 Tom Newby Photography. CC-BY-2.0<\/span><\/div>\n<p>While playing <a href=\"http:\/\/en.wikipedia.org\/wiki\/Articulate!\">Articulate<\/a> with my parents over Christmas, I had to describe &#8216;immense&#8217;, and tried to do so by saying that you could say something was &#8216;of <i>&lt;blank&gt;<\/i> proportions&#8217;. This didn&#8217;t work, but it got me thinking about what adjectives are typically only used to describe one or two things, and, conversely, which nouns are typically described by only one or two adjectives.<\/p>\n<p>Time to dust off <a href=\"http:\/\/googleresearch.blogspot.ie\/2006\/08\/all-our-n-gram-are-belong-to-you.html\">the Google N-Gram data<\/a> from a few years ago, and combine it with <a href=\"http:\/\/wordlist.sourceforge.net\/\">a parts-of-speech database<\/a>.<\/p>\n<h2>Most cliched adjectives<\/h2>\n<p>For each adjective, I found the noun it was most commonly used in front of, and the percentage of uses of that adjective explained by a use before that noun. The adjectives with the highest such percentage are the &#8216;most cliched&#8217;.<\/p>\n<p>Taking the twenty most cliched adjectives, after manually weeding out not-really adjectives, geographical phrases (like &#8216;Saudi Arabia&#8217;), I found the most cliched adjective is:<\/p>\n<table class=\"cliches-results\">\n<tbody>\n<tr>\n<td>91\u00b75%<\/td>\n<td>of the time<\/td>\n<td><b>stainless<\/b><\/td>\n<td>is used, it&#8217;s to describe<\/td>\n<td><b>steel<\/b><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>(This is an approximate statement. A more accurate one would be: 91\u00b75% of the time that &#8216;stainless&#8217; precedes a noun, that noun is &#8216;steel&#8217;. But the above gets the point across.)<\/p>\n<p>The full top twenty list is:<\/p>\n<table class=\"cliches-results\">\n<tbody>\n<tr>\n<td>91\u00b75%<\/td>\n<td><b>stainless<\/b><\/td>\n<td><b>steel<\/b><\/td>\n<td class=\"r-col-grp\">76\u00b75%<\/td>\n<td><b>sporting<\/b><\/td>\n<td><b>goods<\/b><\/td>\n<\/tr>\n<tr>\n<td>89\u00b74%<\/td>\n<td><b>objectionable<\/b><\/td>\n<td><b>content<\/b><\/td>\n<td class=\"r-col-grp\">75\u00b79%<\/td>\n<td><b>stained<\/b><\/td>\n<td><b>glass<\/b><\/td>\n<\/tr>\n<tr>\n<td>87\u00b70%<\/td>\n<td><b>wrought<\/b><\/td>\n<td><b>iron<\/b><\/td>\n<td class=\"r-col-grp\">75\u00b71%<\/td>\n<td><b>motley<\/b><\/td>\n<td><b>fool<\/b><\/td>\n<\/tr>\n<tr>\n<td>84\u00b77%<\/td>\n<td><b>typographical<\/b><\/td>\n<td><b>errors<\/b><\/td>\n<td class=\"r-col-grp\">74\u00b78%<\/td>\n<td><b>designated<\/b><\/td>\n<td><b>trademarks<\/b><\/td>\n<\/tr>\n<tr>\n<td>84\u00b70%<\/td>\n<td><b>elapsed<\/b><\/td>\n<td><b>time<\/b><\/td>\n<td class=\"r-col-grp\">74\u00b75%<\/td>\n<td><b>Grateful<\/b><\/td>\n<td><b>Dead<\/b><\/td>\n<\/tr>\n<tr>\n<td>82\u00b76%<\/td>\n<td><b>martial<\/b><\/td>\n<td><b>arts<\/b><\/td>\n<td class=\"r-col-grp\">73\u00b71%<\/td>\n<td><b>respective<\/b><\/td>\n<td><b>owners<\/b><\/td>\n<\/tr>\n<tr>\n<td>81\u00b79%<\/td>\n<td><b>supreme<\/b><\/td>\n<td><b>court<\/b><\/td>\n<td class=\"r-col-grp\">72\u00b76%<\/td>\n<td><b>vice<\/b><\/td>\n<td><b>president<\/b><\/td>\n<\/tr>\n<tr>\n<td>81\u00b71%<\/td>\n<td><b>movable<\/b><\/td>\n<td><b>type<\/b><\/td>\n<td class=\"r-col-grp\">72\u00b71%<\/td>\n<td><b>deviant<\/b><\/td>\n<td><b>comments<\/b><\/td>\n<\/tr>\n<tr>\n<td>79\u00b73%<\/td>\n<td><b>untitled<\/b><\/td>\n<td><b>document<\/b><\/td>\n<td class=\"r-col-grp\">71\u00b73%<\/td>\n<td><b>Looney<\/b><\/td>\n<td><b>Tunes<\/b><\/td>\n<\/tr>\n<tr>\n<td>76\u00b79%<\/td>\n<td><b>breaking<\/b><\/td>\n<td><b>news<\/b><\/td>\n<td class=\"r-col-grp\">70\u00b79%<\/td>\n<td><b>nervous<\/b><\/td>\n<td><b>system<\/b><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>&#8216;Untitled document&#8217; is pretty good, as is &#8216;Grateful Dead&#8217; and &#8216;Looney Tunes&#8217;.<\/p>\n<h2>Most cliched nouns<\/h2>\n<p>In a similar way, I examined each noun and found the adjective most commonly used to describe it, and the percentage of occurrences of that noun which were paired with that adjective. With this measure, the most cliched noun is:<\/p>\n<table class=\"cliches-results\">\n<tbody>\n<tr>\n<td>97\u00b74%<\/td>\n<td>of the time<\/td>\n<td><b>annotation<\/b><\/td>\n<td>is used, it&#8217;s described as<\/td>\n<td><b>functional<\/b><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>(Again, this is an approximate description. A more accurate one would be: 97\u00b74% of the time &#8216;annotation&#8217; follows an adjective, that adjective is &#8216;functional&#8217;.)<\/p>\n<p>Giving the list as <i>&lt;adjective&gt;<\/i> <i>&lt;noun&gt;<\/i>, the full top twenty list is:<\/p>\n<table class=\"cliches-results\">\n<tbody>\n<tr>\n<td>97\u00b74%<\/td>\n<td><b>functional<\/b><\/td>\n<td><b>annotation<\/b><\/td>\n<td class=\"r-col-grp\">87\u00b77%<\/td>\n<td><b>other<\/b><\/td>\n<td><b>shoppers<\/b><\/td>\n<\/tr>\n<tr>\n<td>97\u00b70%<\/td>\n<td><b>real<\/b><\/td>\n<td><b>estate<\/b><\/td>\n<td class=\"r-col-grp\">86\u00b75%<\/td>\n<td><b>multiple<\/b><\/td>\n<td><b>sclerosis<\/b><\/td>\n<\/tr>\n<tr>\n<td>96\u00b78%<\/td>\n<td><b>creative<\/b><\/td>\n<td><b>commons<\/b><\/td>\n<td class=\"r-col-grp\">86\u00b73%<\/td>\n<td><b>hot<\/b><\/td>\n<td><b>tub<\/b><\/td>\n<\/tr>\n<tr>\n<td>91\u00b76%<\/td>\n<td><b>global<\/b><\/td>\n<td><b>warming<\/b><\/td>\n<td class=\"r-col-grp\">85\u00b76%<\/td>\n<td><b>simple<\/b><\/td>\n<td><b>syndication<\/b><\/td>\n<\/tr>\n<tr>\n<td>91\u00b74%<\/td>\n<td><b>registered<\/b><\/td>\n<td><b>trademark<\/b><\/td>\n<td class=\"r-col-grp\">85\u00b75%<\/td>\n<td><b>due<\/b><\/td>\n<td><b>diligence<\/b><\/td>\n<\/tr>\n<tr>\n<td>91\u00b70%<\/td>\n<td><b>super<\/b><\/td>\n<td><b>saver<\/b><\/td>\n<td class=\"r-col-grp\">84\u00b77%<\/td>\n<td><b>grand<\/b><\/td>\n<td><b>theft<\/b><\/td>\n<\/tr>\n<tr>\n<td>90\u00b78%<\/td>\n<td><b>free<\/b><\/td>\n<td><b>counters<\/b><\/td>\n<td class=\"r-col-grp\">84\u00b72%<\/td>\n<td><b>Iron<\/b><\/td>\n<td><b>Maiden<\/b><\/td>\n<\/tr>\n<tr>\n<td>89\u00b71%<\/td>\n<td><b>planned<\/b><\/td>\n<td><b>parenthood<\/b><\/td>\n<td class=\"r-col-grp\">82\u00b78%<\/td>\n<td><b>remote<\/b><\/td>\n<td><b>sensing<\/b><\/td>\n<\/tr>\n<tr>\n<td>88\u00b75%<\/td>\n<td><b>used<\/b><\/td>\n<td><b>textbooks<\/b><\/td>\n<td class=\"r-col-grp\">82\u00b73%<\/td>\n<td><b>Black<\/b><\/td>\n<td><b>Sabbath<\/b><\/td>\n<\/tr>\n<tr>\n<td>88\u00b71%<\/td>\n<td><b>national<\/b><\/td>\n<td><b>aeronautics<\/b><\/td>\n<td class=\"r-col-grp\">81\u00b77%<\/td>\n<td><b>self<\/b><\/td>\n<td><b>catering<\/b><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Funny to see two British metal bands in there. I think the presence of &#8216;super saver&#8217; and &#8216;other shoppers&#8217; might indicate a large contribution to the corpus from commerce sites.<\/p>\n<h2>What about &#8216;immense&#8217;?<\/h2>\n<p>But if I had to try to get someone to say &#8216;immense&#8217; by giving a noun it&#8217;s commonly used before, what would that noun be? There are two ways you could answer that. For all nouns, find the fraction of times it&#8217;s preceded by &#8216;immense&#8217; and pick the highest. Or, for all nouns find the rank of &#8216;immense&#8217; in the list of adjectives it&#8217;s preceded by, and choose the highest (i.e., numerically smallest) rank. These aren&#8217;t necessarily the same thing, but it turns out they are, and the best you can do is:<\/p>\n<table class=\"cliches-results\">\n<tbody>\n<tr>\n<td>2\u00b70%<\/td>\n<td>of the time<\/td>\n<td><b>multitude<\/b><\/td>\n<td>is used, it&#8217;s described as<\/td>\n<td><b>immense<\/b><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>And &#8216;immense&#8217; is the sixth-most common adjective used to describe a multitude; the top ten being:<\/p>\n<table class=\"cliches-results\">\n<tbody>\n<tr>\n<td>40\u00b74%<\/td>\n<td><b>great<\/b><\/td>\n<td class=\"r-col-grp\">2\u00b70%<\/td>\n<td><b>immense<\/b><\/td>\n<\/tr>\n<tr>\n<td>11\u00b73%<\/td>\n<td><b>whole<\/b><\/td>\n<td class=\"r-col-grp\">1\u00b73%<\/td>\n<td><b>countless<\/b><\/td>\n<\/tr>\n<tr>\n<td>8\u00b72%<\/td>\n<td><b>vast<\/b><\/td>\n<td class=\"r-col-grp\">1\u00b73%<\/td>\n<td><b>infinite<\/b><\/td>\n<\/tr>\n<tr>\n<td>5\u00b71%<\/td>\n<td><b>mixed<\/b><\/td>\n<td class=\"r-col-grp\">1\u00b70%<\/td>\n<td><b>any<\/b><\/td>\n<\/tr>\n<tr>\n<td>2\u00b76%<\/td>\n<td><b>assembled<\/b><\/td>\n<td class=\"r-col-grp\">1\u00b70%<\/td>\n<td><b>large<\/b><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>An example adjective where &#8216;best rank&#8217; and &#8216;highest fraction&#8217; do not lead to the same noun is &#8216;striking&#8217;:<\/p>\n<blockquote class=\"cliches\"><p><strong>striking contrast<\/strong> \u2014 the highest rank &#8216;striking&#8217; reaches is when describing &#8216;contrast&#8217;, when it&#8217;s the sixth-most common adjective for &#8216;contrast&#8217; and accounts for 2\u00b71% of descriptions of &#8216;contrast&#8217;.<\/p><\/blockquote>\n<blockquote class=\"cliches\"><p><strong>striking similarity<\/strong> \u2014 the highest fraction &#8216;striking&#8217; reaches is when describing &#8216;similarity&#8217;, when it accounts for 2\u00b78% of descriptions of &#8216;similarity&#8217;, and ranks seventh amongst adjectives for &#8216;similarity&#8217;.<\/p><\/blockquote>\n<h2>And &#8216;proportions&#8217;?<\/h2>\n<p>And somebody with knowledge of the Google N-Gram data, when given the clue &#8216;<i>&lt;blank&gt;<\/i> proportions&#8217;?<\/p>\n<table class=\"cliches-results\">\n<tbody>\n<tr>\n<td>6\u00b79%<\/td>\n<td>of the time<\/td>\n<td><b>proportions<\/b><\/td>\n<td>is used, they&#8217;re described as<\/td>\n<td><b>epic<\/b><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>It wasn&#8217;t a very good clue for &#8216;immense&#8217; at all.<\/p>\n<h2>Representativeness of the data<\/h2>\n<p>The data-set shows its origins as a corpus drawn from the web, in that there are several amusing features of the results which sadly are not suitable for a general readership. Even though I don&#8217;t really have a readership of any sort, I therefore omit them. Sigh.<\/p>\n<h2>Most common doubled word<\/h2>\n<p>A couple of years, a related question came up: What is the most common doubled word? Filtering for unsuitable content, the winners turned out to be:<\/p>\n<table class=\"cliches-results\">\n<tbody>\n<tr>\n<td>1.<\/td>\n<td><b>blah blah<\/b><\/td>\n<td class=\"r-col-grp\">7.<\/td>\n<td><b>no no<\/b><\/td>\n<\/tr>\n<tr>\n<td>2.<\/td>\n<td><b>had had<\/b><\/td>\n<td class=\"r-col-grp\">8.<\/td>\n<td><b>really really<\/b><\/td>\n<\/tr>\n<tr>\n<td>3.<\/td>\n<td><b>very very<\/b><\/td>\n<td class=\"r-col-grp\">9.<\/td>\n<td><b>much much<\/b><\/td>\n<\/tr>\n<tr>\n<td>4.<\/td>\n<td><b>ha ha<\/b><\/td>\n<td class=\"r-col-grp\">10.<\/td>\n<td><b>long long<\/b><\/td>\n<\/tr>\n<tr>\n<td>5.<\/td>\n<td><b>la la<\/b><\/td>\n<td class=\"r-col-grp\">11.<\/td>\n<td><b>Duran Duran<\/b><\/td>\n<\/tr>\n<tr>\n<td>6.<\/td>\n<td><b>big big<\/b><\/td>\n<td class=\"r-col-grp\">12.<\/td>\n<td><b>etc etc<\/b><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>A good showing from band names altogether in that data-set.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Which adjectives are most commonly paired with only a few nouns?  Which nouns are most commonly paired with only a few adjectives?  Combining the Google N-Gram data with a parts-of-speech database yields some answers.<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-311","post","type-post","status-publish","format-standard","hentry","category-uncategorized","comments-off"],"_links":{"self":[{"href":"https:\/\/redfrontdoor.org\/blog\/index.php?rest_route=\/wp\/v2\/posts\/311","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/redfrontdoor.org\/blog\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/redfrontdoor.org\/blog\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/redfrontdoor.org\/blog\/index.php?rest_route=\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/redfrontdoor.org\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=311"}],"version-history":[{"count":3,"href":"https:\/\/redfrontdoor.org\/blog\/index.php?rest_route=\/wp\/v2\/posts\/311\/revisions"}],"predecessor-version":[{"id":3747,"href":"https:\/\/redfrontdoor.org\/blog\/index.php?rest_route=\/wp\/v2\/posts\/311\/revisions\/3747"}],"wp:attachment":[{"href":"https:\/\/redfrontdoor.org\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=311"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/redfrontdoor.org\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=311"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/redfrontdoor.org\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=311"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}