TF-IDF Embeddings and Cosine Search
Goal
You compute TF-IDF embeddings directly with numpy, search documents with cosine similarity, and confirm the price of dimensionality reduction in numbers.
Why it matters
When you call an embedding API, a vector comes out. You can use it without knowing what happens inside, but when search results look strange you cannot narrow down the cause.
TF-IDF is much simpler than a neural embedding, but the core structure is the same. The skeleton is identical: turn text into a vector, normalize the length, and measure similarity with a dot product. So if you leave out normalization here, you can see for yourself how a long document gets hit by every query, and that experience is used as is when you debug real vector search.
The hashing trick in the last step is meaningful too. Reducing dimensions gains you memory and speed, but loses information because of collisions. If you measure in numbers how large that loss actually is, you develop a feel for choosing a vector dimension.
Steps
The working directory is /root/llm. The target is the 30 documents in the docs table.
- Tokenize each document and save it to
/root/llm/doc_tokens.tsv. The format isdocs.id<탭>공백으로 구분한 토큰들(docs.id, tab, tokens separated by spaces), and it is 30 lines. The tokenization rule is lowercase the text, then extract with the regular expression[가-힣a-z0-9]+. - Save the document frequency per term to
/root/llm/df.tsvin the format항<탭>df(term, tab, df). Terms are sorted ascending, and a term that appears several times in one document counts as 1. - Save the IDF to
/root/llm/idf.tsvin the format항<탭>idf(term, tab, idf). The term order is the same as in step 2 and values have six decimal places. The formula isln((1 + N) / (1 + df)) + 1, where N is the number of documents. - Save the TF-IDF matrix to
/root/llm/tfidf.npy. The shape is (number of documents, number of terms) and the term order is the same as in step 2. The values are the occurrence count multiplied by idf, then each row normalized to an L2 length of 1. - Save each document's nearest neighbor to
/root/llm/sim_top.tsvasdocs.id<탭>가장 비슷한 docs.id<탭>점수(docs.id, tab, the most similar docs.id, tab, score). Exclude the document itself, the score has six decimal places, and on a tie pick the one with the smallerdocs.id. - For the 3 queries below, save the top 3 documents to
/root/llm/query_top3.tsvas질의번호<탭>순위<탭>docs.id<탭>점수(query number, tab, rank, tab, docs.id, tab, score). That is 9 lines in total.- Query 1:
인덱스가 왜 안 타는지 알고 싶다(a Korean sentence meaning "I want to know why the index is not used") - Query 2:
TCP 연결이 안 될 때 무엇을 보나(a Korean sentence meaning "what do you look at when a TCP connection fails") - Query 3:
GPU 는 왜 메모리 때문에 느려지나(a Korean sentence meaning "why does a GPU slow down because of memory") The query vector uses the same tokenization and the same idf, and is L2-normalized. On a tie, the one with the smallerdocs.idcomes first.
- Query 1:
- Build a 1024-dimensional matrix with the hashing trick and save it to
/root/llm/hashed.npy. The bucket isint(md5(항).hexdigest(), 16) % 1024(here the placeholder stands for the term), and after accumulating등장 횟수 곱하기 idf(occurrence count times idf) into each bucket, L2-normalize the rows. - Save the number and the ratio of documents whose nearest neighbors agree between the two methods to
/root/llm/compare.tsvas one line in the format일치수<탭>비율(agreement count, tab, ratio). The ratio has three decimal places.
Notes
numpyis installed. You usenp.save,np.loadandnp.linalg.norm.- Once you fix the order of documents and terms, you must keep the same order to the end.
- Common mistake 1: Python's built-in
hash()gives different values from run to run. Be sure to usehashlib.md5. - Common mistake 2: if you leave out normalization, a long document gets hit by every query. Check that the length of each row is 1.
Tokenize the documents
Tokenize each document and save it to /root/llm/doc_tokens.tsv. The format is docs.id<탭>공백으로 구분한 토큰들 (docs.id, tab, tokens separated by spaces), and it is 30 lines. The tokenization rule is lowercase the text, then extract with the regular expression [가-힣a-z0-9]+.
After lowercasing, keep only the runs of Hangul, English letters and digits. A single regular expression is enough.
Count document frequencies
Save the document frequency per term to /root/llm/df.tsv in the format 항<탭>df (term, tab, df). Terms are sorted ascending, and a term that appears several times in one document counts as 1.
A term that appears several times in one document counts as 1. Sort the terms in ascending order.
Calculate the IDF
Save the IDF to /root/llm/idf.tsv in the format 항<탭>idf (term, tab, idf). The term order is the same as in step 2 and values have six decimal places. The formula is ln((1 + N) / (1 + df)) + 1, where N is the number of documents.
Use the smoothed formula that adds 1 to the denominator and the numerator and adds 1 at the end.
Build the TF-IDF matrix
Save the TF-IDF matrix to /root/llm/tfidf.npy. The shape is (number of documents, number of terms) and the term order is the same as in step 2. The values are the occurrence count multiplied by idf, then each row normalized to an L2 length of 1.
Multiply the occurrence count by idf, then set the length to 1 row by row. If you change the order, the values change.
Find the nearest neighbors
Save each document's nearest neighbor to /root/llm/sim_top.tsv as docs.id<탭>가장 비슷한 docs.id<탭>점수 (docs.id, tab, the most similar docs.id, tab, score). Exclude the document itself, the score has six decimal places, and on a tie pick the one with the smaller docs.id.
Because the vectors are normalized, the dot product is exactly the cosine similarity. Remove the document itself from the candidates.
Search documents with a query
For the 3 queries below, save the top 3 documents to /root/llm/query_top3.tsv as 질의번호<탭>순위<탭>docs.id<탭>점수 (query number, tab, rank, tab, docs.id, tab, score). That is 9 lines in total.
- Query 1:
인덱스가 왜 안 타는지 알고 싶다(a Korean sentence meaning "I want to know why the index is not used") - Query 2:
TCP 연결이 안 될 때 무엇을 보나(a Korean sentence meaning "what do you look at when a TCP connection fails") - Query 3:
GPU 는 왜 메모리 때문에 느려지나(a Korean sentence meaning "why does a GPU slow down because of memory") The query vector uses the same tokenization and the same idf, and is L2-normalized. On a tie, the one with the smallerdocs.idcomes first.
Build the query vector with the same idf and normalize it. Ignore words that were not in the training data.
Reduce dimensions with the hashing trick
Build a 1024-dimensional matrix with the hashing trick and save it to /root/llm/hashed.npy. The bucket is int(md5(항).hexdigest(), 16) % 1024 (here the placeholder stands for the term), and after accumulating 등장 횟수 곱하기 idf (occurrence count times idf) into each bucket, L2-normalize the rows.
Python's built-in hash gives different values from run to run. You have to use md5 to make it reproducible.
Compare the results of the two methods
Save the number and the ratio of documents whose nearest neighbors agree between the two methods to /root/llm/compare.tsv as one line in the format 일치수<탭>비율 (agreement count, tab, ratio). The ratio has three decimal places.
Count whether each document's nearest neighbor is the same in the two methods, and compute the ratio.