Articles in this section

Word Familiarity

What it does

This KPI evaluates the familiarity of each word in a marketing asset, using a comprehensive database of word frequencies across languages. It assesses how easy the text is for the target audience to process, with less familiar words typically requiring more effort to understand.

 

Why it matters

Word familiarity directly influences audience comprehension and engagement. Familiar language ensures the message is clear, relatable, and easy to understand, fostering a stronger connection with the audience and enhancing the effectiveness of communication.

 

How it works

Words are extracted from assets using a sophisticated text detection AI model. Each word's familiarity is then gauged using a vast word frequency database, with the overall score representing the average familiarity of all words in the text, excluding brand names.

 

How to improve

  • Utilize language that resonates with your consumer base, avoiding overly complex or technical terms.
  • Steer clear of industry jargon, except where it adds value to specialized audiences.
  • When technical terms are necessary, balance them with simpler language to maintain accessibility.

 

AI models used

An advanced AI model extracts text from assets, while familiarity is based on word frequency in a huge text data base with millions of data points. This database, updated regularly, reflects real-world word frequency across multiple languages.

 

Science Background

Word familiarity is the relative ease of perception attributed to every word. For example, the two words “encounter” and “meeting” could be used in a similar way, but “meeting” is cognitively easier than “encounter”. Within the past few decades, attempts have been made to measure word familiarity through human experiments within the psycholinguistic domain (see Tanaka-Ishii and Terada, 2018, for review). Studies have generated several word familiarity lists such as Wilson’s list (Wilson, 1988) and the MRC database (MRC Psycholinguistic Database, 2006) in English, which consists of several thousand words with familiarity ratings. In Japanese, Amano’s list (Amano and Kondo, 2000) contains about 70,000 pairs of words and corresponding ratings.

Word Frequency Effect

The word frequency effect refers to the observation that high-frequency words are processed more efficiently than low-frequency words. Frequency of word occurrence is one of the strongest predictors of processing efficiency. High-frequency words are known to more people and are processed faster than low-frequency words (the word frequency effect; Monsell, Doyle, & Haggard, 1989).

There exist many different readability formulae, some of which were conceived years ago. Their continued popularity now is a testimony to the need for an efficient means to evaluate the difficulty of text. Several formulae are provided in text editing software (eg, in Microsoft Word) or made available in online tools. Even though they are extremely popular, there is little evidence that simplifying text using these formulae is associated with increased understanding. Word length is used as a stand-in for word difficulty and is measured in characters or syllables. However, examples demonstrate that this is not always an accurate indicator of word difficulty: ‘disorientation’ or ‘diabetes’ would be considered more difficult than ‘apnea’ by most formulae, but in many cases people know the meaning of the first words but not the last.

Gundy & Kauchak (2014) evaluated word familiarity rather than word length as a stand-in for word difficulty. Their study is the first study to focus on actual difficulty, measured with a multiple-choice task, in addition to perceived difficulty, measured with a Likert scale. Actual difficulty was correlated with word familiarity (r=0.219, p<0.001) but not with word length (r=−0.075, p=0.107). Perceived difficulty was correlated with both word familiarity (r=−0.397, p<0.001) and word length (r=0.254, p<0.001).

The results show that words with a lower frequency of occurrence are more difficult and less often correctly defined by participants.

Word Familiarity of ~300.000 English words 

High Familiarity Low Familiarity
the elastics
new partage
people durance

 

 

Was this article helpful?
3 out of 3 found this helpful