TT Lab
Get started
Learn Learning paths Courses

LLM Engineering

TF-IDF Embeddings and Cosine Search

Continue in TT Lab

Goal

You compute TF-IDF embeddings directly with numpy, search documents with cosine similarity, and confirm the price of dimensionality reduction in numbers.

Why it matters

When you call an embedding API, a vector comes out. You can use it without knowing what happens inside, but when search results look strange you cannot narrow down the cause.

TF-IDF is much simpler than a neural embedding, but the core structure is the same. The skeleton is identical: turn text into a vector, normalize the length, and measure similarity with a dot product. So if you leave out normalization here, you can see for yourself how a long document gets hit by every query, and that experience is used as is when you debug real vector search.

The hashing trick in the last step is meaningful too. Reducing dimensions gains you memory and speed, but loses information because of collisions. If you measure in numbers how large that loss actually is, you develop a feel for choosing a vector dimension.

Steps

The working directory is /root/llm. The target is the 30 documents in the docs table.

  1. Tokenize each document and save it to /root/llm/doc_tokens.tsv. The format is docs.id<탭>공백으로 구분한 토큰들 (docs.id, tab, tokens separated by spaces), and it is 30 lines. The tokenization rule is lowercase the text, then extract with the regular expression [가-힣a-z0-9]+.
  2. Save the document frequency per term to /root/llm/df.tsv in the format 항<탭>df (term, tab, df). Terms are sorted ascending, and a term that appears several times in one document counts as 1.
  3. Save the IDF to /root/llm/idf.tsv in the format 항<탭>idf (term, tab, idf). The term order is the same as in step 2 and values have six decimal places. The formula is ln((1 + N) / (1 + df)) + 1, where N is the number of documents.
  4. Save the TF-IDF matrix to /root/llm/tfidf.npy. The shape is (number of documents, number of terms) and the term order is the same as in step 2. The values are the occurrence count multiplied by idf, then each row normalized to an L2 length of 1.
  5. Save each document's nearest neighbor to /root/llm/sim_top.tsv as docs.id<탭>가장 비슷한 docs.id<탭>점수 (docs.id, tab, the most similar docs.id, tab, score). Exclude the document itself, the score has six decimal places, and on a tie pick the one with the smaller docs.id.
  6. For the 3 queries below, save the top 3 documents to /root/llm/query_top3.tsv as 질의번호<탭>순위<탭>docs.id<탭>점수 (query number, tab, rank, tab, docs.id, tab, score). That is 9 lines in total.
    • Query 1: 인덱스가 왜 안 타는지 알고 싶다 (a Korean sentence meaning "I want to know why the index is not used")
    • Query 2: TCP 연결이 안 될 때 무엇을 보나 (a Korean sentence meaning "what do you look at when a TCP connection fails")
    • Query 3: GPU 는 왜 메모리 때문에 느려지나 (a Korean sentence meaning "why does a GPU slow down because of memory") The query vector uses the same tokenization and the same idf, and is L2-normalized. On a tie, the one with the smaller docs.id comes first.
  7. Build a 1024-dimensional matrix with the hashing trick and save it to /root/llm/hashed.npy. The bucket is int(md5(항).hexdigest(), 16) % 1024 (here the placeholder stands for the term), and after accumulating 등장 횟수 곱하기 idf (occurrence count times idf) into each bucket, L2-normalize the rows.
  8. Save the number and the ratio of documents whose nearest neighbors agree between the two methods to /root/llm/compare.tsv as one line in the format 일치수<탭>비율 (agreement count, tab, ratio). The ratio has three decimal places.

Notes

Tokenize the documents

Tokenize each document and save it to /root/llm/doc_tokens.tsv. The format is docs.id<탭>공백으로 구분한 토큰들 (docs.id, tab, tokens separated by spaces), and it is 30 lines. The tokenization rule is lowercase the text, then extract with the regular expression [가-힣a-z0-9]+.

After lowercasing, keep only the runs of Hangul, English letters and digits. A single regular expression is enough.

Count document frequencies

Save the document frequency per term to /root/llm/df.tsv in the format 항<탭>df (term, tab, df). Terms are sorted ascending, and a term that appears several times in one document counts as 1.

A term that appears several times in one document counts as 1. Sort the terms in ascending order.

Calculate the IDF

Save the IDF to /root/llm/idf.tsv in the format 항<탭>idf (term, tab, idf). The term order is the same as in step 2 and values have six decimal places. The formula is ln((1 + N) / (1 + df)) + 1, where N is the number of documents.

Use the smoothed formula that adds 1 to the denominator and the numerator and adds 1 at the end.

Build the TF-IDF matrix

Save the TF-IDF matrix to /root/llm/tfidf.npy. The shape is (number of documents, number of terms) and the term order is the same as in step 2. The values are the occurrence count multiplied by idf, then each row normalized to an L2 length of 1.

Multiply the occurrence count by idf, then set the length to 1 row by row. If you change the order, the values change.

Find the nearest neighbors

Save each document's nearest neighbor to /root/llm/sim_top.tsv as docs.id<탭>가장 비슷한 docs.id<탭>점수 (docs.id, tab, the most similar docs.id, tab, score). Exclude the document itself, the score has six decimal places, and on a tie pick the one with the smaller docs.id.

Because the vectors are normalized, the dot product is exactly the cosine similarity. Remove the document itself from the candidates.

Search documents with a query

For the 3 queries below, save the top 3 documents to /root/llm/query_top3.tsv as 질의번호<탭>순위<탭>docs.id<탭>점수 (query number, tab, rank, tab, docs.id, tab, score). That is 9 lines in total.

Build the query vector with the same idf and normalize it. Ignore words that were not in the training data.

Reduce dimensions with the hashing trick

Build a 1024-dimensional matrix with the hashing trick and save it to /root/llm/hashed.npy. The bucket is int(md5(항).hexdigest(), 16) % 1024 (here the placeholder stands for the term), and after accumulating 등장 횟수 곱하기 idf (occurrence count times idf) into each bucket, L2-normalize the rows.

Python's built-in hash gives different values from run to run. You have to use md5 to make it reproducible.

Compare the results of the two methods

Save the number and the ratio of documents whose nearest neighbors agree between the two methods to /root/llm/compare.tsv as one line in the format 일치수<탭>비율 (agreement count, tab, ratio). The ratio has three decimal places.

Count whether each document's nearest neighbor is the same in the two methods, and compute the ratio.