Blog
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 3525 posts
#2026-03 765#english 592#culture 264#deep-dive 254#kubernetes 249#career 229#ai 219#llm 209#devops 196#2026-04 146#security 143#database 114#observability 113#communication 109#history 107#architecture 100#productivity 96#finance 88#economy 84#mindset 81#psychology 80#ai-papers 79#food 78#it 78#travel 78#deep-learning 77#japanese 77#networking 77#performance 72#business-travel 70#linux 70#gpu 69#ai-agent 66#cs-fundamentals 63#postgresql 61#rag 58#self-improvement 55#learning 53#mlops 53#python 52
Anatomy of config.json — Reading a Model From One Settings File
How every field in config.json — hiddensize, numhiddenlayers, numattentionheads versus numkeyvalueheads, headdim, intermediatesize, ropetheta, vocabsize, tiewordembeddings — shows up in memory and speed, and a hand count
2026-08-12 · 6 min read #ai-papers#model-internals#config-json#transformer#llm-architectureOCR and Document Understanding Technical Reports: What to Read, and Why Parsing Is Not Finished
Ten OCR and document understanding technical reports, each verified by opening the arXiv abstract page directly. From Donut and Nougat through GOT-OCR2.0, olmOCR, DeepSeek-OCR and its successor, GLM-OCR, Qianfan-OCR, Min
2026-08-12 · 7 min read #ai-papers#paper-review#technical-report#ocr#document-aiHow to Read Leaderboards and Benchmarks: Why SOTA Has Such a Short Shelf Life
Twelve benchmark methodology papers, each verified by opening the arXiv abstract page directly, assembled into a guide for reading leaderboard numbers. Data contamination, prompt format sensitivity, eval harness differen
2026-08-12 · 7 min read #ai-papers#paper-review#technical-report#benchmark#evaluationReading a Technical Report Critically — What Gets Written Down and What Goes Missing
How to separate the verifiable from the unverifiable in a model technical report. Covers the limits of self-reported benchmarks, the places where a config and a report disagree, values that could not be read from gated r
2026-08-12 · 7 min read #ai-papers#model-internals#tech-report#benchmarks#evaluationMoE Routing — How an Expert Gets Picked
Reading the config fields of a mixture-of-experts layer against real models. Comparing Mixtral 2-of-8, Qwen3 8-of-128, the 256 routed experts plus a shared expert in DeepSeek-V3, and the sparsity of 48 in Kimi K2, then w
2026-08-12 · 6 min read #ai-papers#model-internals#mixture-of-experts#moe#routingSpeech Recognition and Synthesis Technical Reports: What to Read, and What a Single WER Hides
Ten speech recognition and synthesis technical reports, each verified by opening the arXiv abstract page directly. Whisper, Omnilingual ASR, Qwen3-ASR, the Open ASR Leaderboard, Seed-TTS and F5-TTS, CosyVoice 2, Qwen3-TT
2026-08-12 · 7 min read #ai-papers#paper-review#technical-report#speech-recognition#text-to-speechImage Generation and Understanding Technical Reports: What to Read, and Why Design Beats Sample Images
Eleven image generation and understanding technical reports, each verified by opening the arXiv abstract page directly. Rectified flow transformers, VAR, Emu3, SANA, Janus-Pro, FLUX.1 Kontext, the Qwen-Image line, Seedre
2026-08-12 · 8 min read #ai-papers#paper-review#technical-report#image-generation#diffusion-transformerInside the Tokenizer — Why Korean Costs More Tokens, and What That Costs
How byte-level BPE works, then downloading the actual tokenizer files of Qwen3, DeepSeek-V3, and Mixtral to tokenize the same English and Korean text and compare. Covers the vocabulary-size tradeoff, why the config vocab
2026-08-12 · 6 min read #ai-papers#model-internals#tokenizer#bpe#korean-nlpNormalization and Activation — Keeping Training From Falling Apart
Confirming from config values why RMSNorm, pre-norm, and SwiGLU became the defaults, then walking through the newer devices that stop attention logits from exploding — the QK-Norm of Qwen3 and the QK-Clip of Kimi K2 — wi
2026-08-12 · 6 min read #ai-papers#model-internals#rmsnorm#swiglu#training-stabilityAttention Variants — From MHA to MLA, and How the KV Cache Shrinks
Comparing MHA, MQA, GQA, and MLA using real config values. Working out with formulas and numbers how the grouped-query attention in Qwen3, Mixtral, and GLM-4.5 and the latent attention in DeepSeek-V3 and Kimi K2 reduce p
2026-08-12 · 6 min read #ai-papers#model-internals#attention#gqa#mlaWhat Is Not in the Asset Inventory Never Gets Scanned — OT Exposure Management, From the Water Utility PLC Case
On 30 July 2026 CISA warned of a sharp rise in activity targeting PLCs in the water and wastewater sector and urged operators to remove internet-exposed OT immediately. The behavior the advisory observed was not compromi
2026-08-09 · 9 min read #security#ot#ics#plc#cisaA Breach With No Attacker — Why Agent Credentials Deserve Another Look
Hugging Face disclosed a production breach caused by autonomous agents on 16 July 2026, and about three weeks later OpenAI revealed that the attack had leaked out of its own training environment. This post is not an inci
2026-08-09 · 8 min read #security#llm#agent#incident-response#credentialsA Single Instruction Can Take 62 Seconds — Latency Is a Property of the Path, Not of the Instruction
The Assembly Hall of Shame is a leaderboard for the competition to make a single instruction as slow as possible. At the bottom, nop takes 1 cycle; at the top, fxrstor64 takes 198 billion cycles, or 62 seconds. Read that
2026-08-09 · 11 min read #os-concepts#performance#cpu#microarchitecture#benchmarkWhy Diátaxis Gets Mistaken for Four Folders — And Why Two Modes in One Page Collapse
Diátaxis divides documentation into four kinds: tutorial, how-to, reference, and explanation. Yet most teams finish their adoption by creating four folders, and the actual problem stays exactly where it was. The real fai
2026-08-09 · 9 min read #documentation#diataxis#technical-writing#developer-experience#information-architectureHow a Domain Says It Is For Sale — When DNS Became a Channel for Claims
A domain can be registered, serving a perfectly healthy site, and still be for sale — and until now a machine had no way to find that out. RFC 10023 defines that signal with a single underscore-prefixed node name and a T
2026-08-09 · 7 min read #dns#rfc#networking#domain#protocolThe Exact Scope of the Phrase x86 Hardware Backdoor — Reading rosenbridge as Its Author Wrote It
The repository title says hardware backdoors in x86 CPUs, but the body of the README states that the only thing believed to be affected is the VIA C3 and that later generations no longer carry the feature. In the disclai
2026-08-09 · 9 min read #security#hardware#x86#cpu#fuzzingA Harness Is Not Configuration but a Deployable — The Real Bottleneck of the Self-Improvement Loop
The harness engineering post Lilian Weng published in July 2026 gives a name to the whole system wrapped around a model and treats it as a single engineering object. Rather than restating that definition, this post cover
2026-08-09 · 7 min read #llm#agent#harness-engineering#context-engineering#evaluationDoes Running Five Agents Really Make You Five Times Faster — The Bottleneck in Parallel Work Is Not Generation
Purpose-built environments for running several coding agents at once are multiplying. Orca is one of them, and it pitches spraying a single prompt at several agents, isolating each in its own git worktree, then comparing
2026-08-09 · 9 min read #developer-tools#git#worktree#code-review#workflowVisitor Analytics Shows Only 0.5 Percent of Your Traffic — Judge Bots by Origin, Not by Self-Report
Drawing on one year of defending a 1.5-million-page site against scrapers, this post sets out the principles of dealing with bot traffic. JavaScript-based analytics cannot count bots, so you have to read server logs, and
2026-08-09 · 11 min read #network#bot#cloudflare#waf#scrapingWhen Does a Code Screenshot in Your Docs Start Lying — Turning Images into Build Artifacts
A code screenshot captured by hand and pasted into a document is an artifact with no source. The code changes and the image does not, so at some point it quietly starts showing something false. Generate the image from a
2026-08-09 · 8 min read #documentation#developer-tools#cli#ci#automation