Blog
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 3525 posts
#2026-03 765#english 592#culture 264#deep-dive 254#kubernetes 249#career 229#ai 219#llm 209#devops 196#2026-04 146#security 143#database 114#observability 113#communication 109#history 107#architecture 100#productivity 96#finance 88#economy 84#mindset 81#psychology 80#ai-papers 79#food 78#it 78#travel 78#deep-learning 77#japanese 77#networking 77#performance 72#business-travel 70#linux 70#gpu 69#ai-agent 66#cs-fundamentals 63#postgresql 61#rag 58#self-improvement 55#learning 53#mlops 53#python 52
Training Vision LLMs — How to Teach Input and Output
A vision-language model is trained in stages, from alignment pretraining to instruction fine-tuning. We organize what gets taught and how, from the angle of the training pipeline: vision encoder freezing strategy, data c
2026-06-26 · 17 min read #mlops#vision-language-model#multimodal#training#instruction-tuningLLM Inference Serving 2026 — Comparing vLLM, SGLang, and TensorRT-LLM ♪ Listenable
A clear overview of LLM inference serving in 2026. From core principles such as the difference in nature between prefill and decode, continuous batching, and paged KV cache, to a strengths-and-weaknesses comparison of vL
2026-06-26 · 15 min read #llm-serving#vllm#sglang#tensorrt-llm#inferenceMaking Inference Fast — Speculative Decoding and Throughput Optimization
From the fundamental reason LLM decode is slow, to how speculative decoding boosts speed, variants such as Medusa and EAGLE, chunked prefill and prefill/decode disaggregation, the latency versus throughput trade-off, and
2026-06-26 · 13 min read #speculative-decoding#throughput#inference#mlops#latencyServing Multimodal LLMs — The New Challenges Image Input Creates
From how multimodal LLM serving differs from text-only serving, to the added vision-encoder stage, variable visual token counts, prefill cost spikes, the difficulty of multimodal KV cache and batching, latency decomposit
2026-06-26 · 14 min read #mlops#multimodal#llm-serving#vllm#kv-cacheUnderstanding Positional Encoding — From Sine Waves to RoPE
Starting from why Transformers need positional information, this article explains sinusoidal, learned, and relative positional encodings step by step, then RoPE and ALiBi. It connects length extrapolation and context ext
2026-06-26 · 17 min read #llm#positional-encoding#rope#alibi#long-contextDissecting the Transformer — From Attention to KV Cache
A from-scratch breakdown of the Transformer: self-attention, multi-head, positional encoding, the FFN, and residual connections with normalization. It connects tensor shapes and parameter counts, causal masking, encoder/
2026-06-26 · 16 min read #llm#transformer#attention#positional-encoding#kv-cacheThe Evolution of Attention — MQA, GQA, FlashAttention, and Long Context ♪ Listenable
We analyze the memory and compute cost of standard attention, then explain how MQA and GQA shrink the KV cache and how FlashAttention optimizes IO. We compare sliding-window and long-context techniques and trace how all
2026-06-26 · 18 min read #llm#attention#flashattention#gqa#mqaVision LLM Architecture — How an Image Becomes Language ♪ Listenable
A vision-language model processes an image with a vision encoder, then passes it through a projector to produce tokens an LLM can read. From patch embedding to arbitrary-resolution handling, we trace the full path by whi
2026-06-26 · 20 min read #llm#vision-language-model#multimodal#vit#qwen2-vlMultimodal Tokenization and Fusion — Turning Images and Audio Into Tokens
A deep look at how images, audio, and video become tokens and get woven into one sequence with text. We cover patch and VQ image tokenization, discrete-codec audio tokenization, frame sampling, interleaving and separator
2026-06-26 · 14 min read #llm#multimodal#tokenization#vision-language#q-formerThe KV Cache and PagedAttention — Everything About Inference Memory
A deep dive into the KV cache, the single biggest consumer of memory in LLM inference. We cover what the KV cache is and why it eats memory, the memory arithmetic and fragmentation problem, PagedAttention block managemen
2026-06-26 · 12 min read #kv-cache#paged-attention#inference#gpu-memory#quantizationMultimodal AI Training Methods — Many Senses in One Model
A walkthrough of how multimodal AI learns to handle images, text, audio, and video in a single model. We cover modality alignment and contrastive learning, fusion strategies, shared embedding spaces, pretraining and fine
2026-06-26 · 15 min read #ai-papers#multimodal#clip#vision-language#contrastive-learningRust Macros: Declarative vs Procedural (derive)
Rust macros are a metaprogramming tool that generates code at compile time. This post distinguishes pattern-based declarative macros (macrorules!) from procedural macros (derive, attribute, function-like) that receive co
2026-06-26 · 10 min read #rust#macros#metaprogrammingPRs and Commit Messages That Get Merged Fast
PRs that get reviewed quickly and merged fast have things in common. Small, single-purpose PRs, Conventional Commits, the "why" in the commit body rather than just the "what", a self-review first, a PR description with c
2026-06-26 · 8 min read #git#code-review#collaborationBeyond OCR — OCR-free Document Understanding and Unified Models ♪ Listenable
A traditional OCR pipeline splits into detection, recognition, and layout stages, but errors accumulate. We organize the shift in document AI: Donut-style and VLM-based OCR-free document understanding, high-resolution an
2026-06-26 · 20 min read #ai-papers#ocr-free#document-understanding#multimodal#donutCode Review as Communication: Feedback Without Friction
Code review is a technical activity for catching defects, but it is also a conversation between two people. This post covers framing critique at the code rather than the person, using conventional comments to separate ni
2026-06-26 · 14 min read #code-review#communication#engineeringThe HTTP QUERY Method — A New Verb for an Old REST Dilemma
GET cannot carry a request body and POST is semantically wrong for search. This article explores the HTTP QUERY method that resolves this long-standing dilemma, covering its motivation, semantics, and practical adoption.
2026-06-25 · 14 min read #http#query-method#rest#api-design#rfcDNS in Depth — The World of Name Resolution That Free Hosting Revealed ♪ Listenable
DNS is trending again. From recursion and authority to caching and TTL, global Anycast distribution, DNSSEC, and DoH/DoT, we take a deep look at the entire process of name resolution. As free DNS hosting becomes the norm
2026-06-25 · 22 min read #network#dns#anycast#dnssec#dohThe Rise of AI Code Review Tools — What to Delegate and What Humans Should Still See
AI code review tools are spreading fast. We separate the defects AI catches well from the areas humans must still own, and cover CI integration, an adoption checklist, and a critical perspective.
2026-06-25 · 16 min read #devops#code-review#ai-tools#developer-experience#ci-cdHunting Bugs with AI — The Era of Automated Security Research ♪ Listenable
Cases of AI automatically probing APIs at scale to uncover vulnerabilities are on the rise. From how fuzzing, differential analysis, and LLM-assisted triage work, to the asymmetry of attackers also using AI, the implicat
2026-06-25 · 20 min read #devops#security#ai#bug-bounty#fuzzingTechnical Overreach and Scope — Lessons from Carmacks Retrospective ♪ Listenable
Using John Carmacks publicly shared Quake retrospective as a starting point, we draw out the technical ambition and organizational challenges of early game development as general lessons. We cover the value of choosing s
2026-06-25 · 25 min read #culture#engineering#scope#technical-debt#team