Blog
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 3525 posts
#2026-03 765#english 592#culture 264#deep-dive 254#kubernetes 249#career 229#ai 219#llm 209#devops 196#2026-04 146#security 143#database 114#observability 113#communication 109#history 107#architecture 100#productivity 96#finance 88#economy 84#mindset 81#psychology 80#ai-papers 79#food 78#it 78#travel 78#deep-learning 77#japanese 77#networking 77#performance 72#business-travel 70#linux 70#gpu 69#ai-agent 66#cs-fundamentals 63#postgresql 61#rag 58#self-improvement 55#learning 53#mlops 53#python 52
Making vLLM Fast — Configuration, Internals, and Where to Actually Touch the Code ♪ Listenable
A step-by-step walk through improving vLLM performance, starting from everything you can fix without touching code. Covers batching-related arguments, prefix caching, chunked prefill, and quantization choice first, then
2026-08-02 · 21 min read #vllm#llm-inference#paged-attention#benchmark#schedulerWhat We Lose by Delegating — Automation Doesn't Erode Skill Evenly
Nobody feels guilty using a calculator, but most people feel a little uneasy after sending off a document AI drafted for them. That asymmetry is the question behind this piece. It sorts out the territory where offloading
2026-08-02 · 14 min read #humanities#ai#cognition#automation#skillGPU Compiler and Framework Landscape — One Problem, Turning a Graph into a Kernel, a Different Answer at Every Layer ♪ Listenable
This post puts NVCC and PTX, LLVM, MLIR, Triton, torch.compile, XLA, IREE, and TVM on one map. Different names, different owners, but they all solve the same problem: turning a computation graph into an executable kernel
2026-08-02 · 21 min read #gpu#compiler#mlir#triton#pytorchWhat It Really Means to Hand-Tune a GPU Kernel — Making One Transpose Kernel 5x Faster ♪ Listenable
Starting from threads, warps, and the memory hierarchy, this post covers what it actually means to hand-modify a GPU kernel. It explains why occupancy is a symptom rather than a goal, and why most kernels are bound by me
2026-08-02 · 19 min read #cuda#gpu-kernel#nsight-compute#memory-bandwidth#performanceThe Math You Need for Robotics, in Order: And What You Can Safely Put Off ♪ Listenable
An answer to how far you actually need to take your math to build a robot arm. Organized into six branches in order: linear algebra, trigonometry and rotation representations, calculus and multivariable methods, differen
2026-08-02 · 25 min read #robotics#math#electronics#linear-algebra#controlWhat a Robot Arm Is Made Of: Links, Joints, and What Actually Drives Them ♪ Listenable
Buy six servos to build a robot arm and it almost always collapses under its own weight. This post starts from the precise definitions of link, joint, and end effector, explains what degrees of freedom actually count and
2026-08-02 · 29 min read #robotics#hardware#electronics#arduino#mathThree Layers of Writing a Kernel — Comparing CUDA C++, Triton, and CUTLASS on the Same Problem ♪ Listenable
Compares what changes when you approach the same GPU kernel by hand in CUDA C++, tile-by-tile in Python with Triton, or assembled from templates in CUTLASS and CuTe. We actually write a row-wise softmax in both CUDA C++
2026-08-02 · 18 min read #triton#cuda#cutlass#gpu-kernel#compilerKorean Film and Drama Roundup 2025-2026 — The Year Theaters Came Back, the Year OTT Got Reshuffled
A roundup of 18 Korean films and dramas that actually became talked-about from 2025 through the first half of 2026. Covers The King's Warden, which drew 16.91 million admissions, Park Chan-wook's No Other Choice, which p
2026-08-02 · 17 min read #culture#korean-film#k-drama#movie-recommendation#netflixChinese-Language Film and Drama Roundup 2025-2026 — Three Different Report Cards From the Mainland, Hong Kong, and Taiwan ♪ Listenable
A roundup of 17 Chinese-language films and dramas that actually became talked-about from 2025 through the first half of 2026, sorted by mainland China, Hong Kong, and Taiwan. Covers Ne Zha 2, which became the highest-gro
2026-08-02 · 18 min read #culture#chinese-film#hong-kong-cinema#taiwan-cinema#movie-recommendationJapanese Film, Drama, and Anime Roundup 2025-2026 — Two Records That Got Rewritten, and the Shadow They Cast
A roundup of 18 Japanese films, dramas, and anime that actually became talked-about from 2025 through the first half of 2026. Covers the theatrical Demon Slayer: Infinity Castle film that rewrote Japan's all-time box off
2026-08-02 · 17 min read #culture#japanese-film#anime#j-drama#movie-recommendationAmerican Film and TV Roundup 2025-2026 — What the Oscars Picked and What Audiences Picked
A roundup of 17 American films and series that actually became talked-about from 2025 through the first half of 2026. Covers One Battle After Another, winner of the 98th Academy Award for Best Picture; Sinners, which tie
2026-08-02 · 15 min read #culture#american-film#tv-series#movie-recommendation#oscarsWhere AI-Written Posts Fall Apart — Six Failure Modes and a Guardrail for Each
AI-written posts fail while the sentences stay smooth. This post organizes six failure modes — fabricated sources, information stale past the cutoff, paragraphs stretched to repeat the same point, unsupported assertions,
2026-08-02 · 15 min read #ai-writing#content-quality#hallucination#seo#editingMaking Logs Searchable, and Not Going Broke Doing It — Structuring, Mapping Explosions, Retention, and Real Cost
The point where log costs overtake compute costs arrives for most organizations. What delays that point isn't the compression ratio — it's the decision about what becomes a field. This post covers field design for struct
2026-08-02 · 15 min read #observability#logging#opensearch#elasticsearch#costPersuasive Writing — The Structure That Gets Design Docs and RFCs Approved
For anyone writing design docs, RFCs, proposals, or incident follow-up recommendations. This post covers putting the decision you want at the very top instead of the order in which you solved the problem, why showing the
2026-08-02 · 14 min read #career#writing#persuasion#design-doc#engineeringHow to Handle Objections — Telling Factual, Value, and Status Disagreements Apart
Objections are the part most people learning persuasion skip. This post covers steelmanning as a working procedure — restating the other side's argument better than they did, out loud, and getting it confirmed — why rais
2026-08-02 · 14 min read #career#persuasion#communication#conflict#teamworkTaking Apart LLM Benchmark Tooling — Why the Same MMLU Gives Different Scores in Different Harnesses ♪ Listenable
A benchmark score is not a property of the model — it is the result of a measurement performed under specific conditions. The incident where the same LLaMA 65B scored 63.6 and 48.8 on the same MMLU at the same time shows
2026-08-02 · 21 min read #llm-evaluation#benchmark#lm-evaluation-harness#helm#reproducibilityBuilding an Eval Set for Your Own Service — From Traffic Collection to Statistical Significance ♪ Listenable
Public benchmarks cannot measure your problem for you, because the input distribution, the definition of a correct answer, and the cost structure of failure are all different. This post walks through, with real code, how
2026-08-02 · 23 min read #llm-evaluation#eval-set#rubric#statistics#regression-testingHow Text, Images, and Agents Are Each Measured — Why the Three Domains Measure Fundamentally Different Things ♪ Listenable
Text, images, and agents all use the word "performance," but their measurement structures are entirely different. Text splits into multiple-choice that pretends to have a correct answer and open-ended generation that has
2026-08-02 · 21 min read #llm-evaluation#multimodal#llm-as-judge#agent-benchmark#metricsThe Structure of Persuasion — It's Sequence, Not Eloquence, That Moves People
Persuasion fails at the level of structure long before it fails at the level of phrasing. This post is organized around four axes: establishing a shared premise before you make your claim, putting your strongest evidence
2026-08-02 · 17 min read #career#persuasion#communication#influence#psychologyInstrumenting Your App With OpenTelemetry — From Auto-Instrumentation to Manual Spans, and Why You Put a Collector in Front
Instrumentation isn't about installing an SDK — it's about following an order of operations. This post builds a skeleton in a day with auto-instrumentation, locks down resource attributes first, and shows the order for a
2026-08-02 · 16 min read #observability#opentelemetry#instrumentation#otel-collector#tracing