20016 shaares
The two families of similarity
The first family is lexical similarity. This is where local libraries make light work of the problem. If I want to match names, organisations, cities, account references, or noisy OCR output, then normalisation, entity extraction, fuzzy matching, and character n-grams are often fast, cheap, and surprisingly effective.
The second family is semantic similarity. This is where embeddings start to matter. If two phrases are conceptually related but use different wording, fuzzy matching often falls apart. Embeddings give you a vector representation that lets you compare meaning rather than spelling.