Blog
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 3525 posts
#2026-03 765#english 592#culture 264#deep-dive 254#kubernetes 249#career 229#ai 219#llm 209#devops 196#2026-04 146#security 143#database 114#observability 113#communication 109#history 107#architecture 100#productivity 96#finance 88#economy 84#mindset 81#psychology 80#ai-papers 79#food 78#it 78#travel 78#deep-learning 77#japanese 77#networking 77#performance 72#business-travel 70#linux 70#gpu 69#ai-agent 66#cs-fundamentals 63#postgresql 61#rag 58#self-improvement 55#learning 53#mlops 53#python 52
Open Source Worth Watching Right Now (5) Data and ML Pipelines
A data pipeline is not something a single scheduler solves. Ingestion, transformation, orchestration, execution engines, the model lifecycle, and search stores have each become the territory of a different tool. This pos
2026-08-12 · 5 min read #open-source#data-engineering#mlops#python#rustThe Generational Shift in Languages and Runtimes — Eleven Projects Read Through Their Official EOL Notices
Eleven languages, frameworks and runtimes with official end-of-life notices on record. Python 2, AngularJS, Vue 2, Nashorn, Java applets and Web Start, Mono, Xamarin, PhoneGap, Atom, io.js and jQuery Mobile, each split i
2026-08-12 · 10 min read #open-source#javascript#java#python#frameworkBuild Tools That Stepped Down From Default — Why Ten Frontend Toolchains Gave Up Their Place
Ten tools that were once the default in frontend projects and are rarely picked for new ones today, organised strictly around verifiable evidence: official deprecation notices and repository archive status. Create React
2026-08-12 · 9 min read #open-source#frontend#build-tools#deprecation#javascriptGrowing into a Harness Engineer — Why the Job Exists and What to Practice
The title harness engineer is still rare in job postings, but the work already exists in every team shipping agents. The final part 8 of the harness engineering series covers why this job emerged, how existing software s
2026-08-12 · 4 min read #llm#agent#harness-engineering#하네스엔지니어링#AI에이전트Loop Design — Between Infinite Loops and Giving Up Early
Agent loops fail in two directions: the infinite loop that repeats the same call dozens of times, and the early stop that quits at the first obstacle. Part 4 of the harness engineering series covers retry caps, the three
2026-08-12 · 5 min read #llm#agent#harness-engineering#하네스엔지니어링#AI에이전트The Evaluator Bottleneck — A Weak Grader Caps the Whole System
If the score does not move no matter how much you fix the harness, the bottleneck may be the evaluator, not the harness. You cannot select for a quality you cannot measure, which is why a weak grader becomes the ceiling
2026-08-12 · 5 min read #llm#agent#harness-engineering#하네스엔지니어링#AI에이전트Reward Hacking — The Metric Rises While the Task Fails
If the cheapest way for an agent to pass the tests is to edit the tests, the agent will edit the tests. Reward hacking is not a bug; it is the exact optimization of the goal we wrote down. Part 6 of the harness engineeri
2026-08-12 · 5 min read #llm#agent#harness-engineering#하네스엔지니어링#AI에이전트Open Source Worth Watching Right Now (6) What Star Counts Do Not Tell You
A star count is a popularity metric, not a risk metric. This post lays out the signals you actually have to check before you bring an open source project into production: recent commits and release cadence, issue respons
2026-08-12 · 6 min read #open-source#governance#supply-chain#risk#devopsOpen Source Worth Watching Right Now (1) AI Agents and LLM Tooling
The LLM application stack has split into layers: inference servers, orchestration, gateways, agents, and RAG. This post introduces 12 open source projects that are actually used at each layer, grouped by role rather than
2026-08-12 · 6 min read #open-source#llm#ai-agent#ai-platform#ragThe Generational Shift in Container Infrastructure — What Eleven Projects Left Behind
Eleven projects that were once standard parts of a container infrastructure and have since been replaced, documented using only official notices and repository archive status as evidence. rkt, dockershim, Classic Swarm,
2026-08-12 · 10 min read #open-source#kubernetes#container#docker#infrastructureTool Surface Design — One Schema Line Moves the Success Rate
Adding more tools and watching the agent success rate drop is not rare. The tool surface is the agent interface, and the names, descriptions, parameters, failure returns, and response sizes are all design material. Part
2026-08-12 · 5 min read #llm#agent#harness-engineering#하네스엔지니어링#AI에이전트The Context Budget — Design Is What You Leave Out, Not What You Put In
The context window still has room, yet agent accuracy is dropping. Context is a finite attention budget, and tool schemas spend it too. Part 2 of the harness engineering series covers turning prompt accumulation into a p
2026-08-12 · 6 min read #llm#agent#harness-engineering#하네스엔지니어링#AI에이전트Harness Fingerprints and Versioning — Making Unrecorded Changes Traceable
The success rate moved with no prompt commit and no model change — so what do you roll back? Part 7 of the harness engineering series covers the harness fingerprint: a single normalized hash summarizing every decision th
2026-08-12 · 5 min read #llm#agent#harness-engineering#하네스엔지니어링#AI에이전트What Is Harness Engineering — The Model Is a Fixed Input; What You Ship Is Everything Around It
Two teams use the same model, so why do their agents perform so differently? For most teams the model is a fixed input, and what actually ships is the harness around it: the tool surface, the failure return format, the l
2026-08-12 · 6 min read #llm#agent#harness-engineering#하네스엔지니어링#AI에이전트vLLM Metrics — What to Chart and What to Alert On
The series vLLM exposes answer questions GPU metrics cannot: how many requests are running versus waiting right now, how full the KV cache is, how long until the first token. This post reads the official vLLM documentati
2026-08-12 · 7 min read #gpu#kubernetes#vllm#prometheus#observabilityMIG and Time-Slicing — Two Ways to Share One GPU
There are broadly two ways to put multiple workloads on one GPU: time-slicing, which divides time, and MIG, which divides hardware. Despite sounding similar, their isolation guarantees are nothing alike. This post works
2026-08-12 · 6 min read #gpu#kubernetes#mig#time-slicing#nvidiaDCGM Exporter — GPU Utilization Is Not What You Think It Is
DCGM Exporter is the standard path for exposing GPU telemetry in Prometheus format, but the utilization-style metric that ends up on nearly every dashboard does not measure what people expect. This post reads the default
2026-08-12 · 9 min read #gpu#kubernetes#dcgm#prometheus#observabilityA GPU Troubleshooting Playbook — Fix the Layers, Then Walk Down
The biggest waste in diagnosing GPU problems on Kubernetes is poking around without an order. A pod that will not schedule, a pod that runs but cannot see the GPU, a driver and toolkit version mismatch, memory exhaustion
2026-08-12 · 7 min read #gpu#kubernetes#troubleshooting#nvidia#gpu-operatorNVIDIA GPU Operator — The Six Pieces You Used to Install by Hand
Running GPUs on Kubernetes used to mean matching six pieces on every node by hand: the driver, the NVIDIA Container Toolkit, the device plugin, DCGM, GPU Feature Discovery, and Node Feature Discovery. The NVIDIA GPU Oper
2026-08-12 · 6 min read #gpu#kubernetes#gpu-operator#nvidia#dcgmKorean Dev Blog Curation 2 — Incident Retrospectives and Troubleshooting, 12 Posts I Opened and Checked
The genre Korean developers write best is the incident retrospective. Twelve posts: p6spy silently defeating read/write datasource routing, why an HTTP timeout does not cover DNS resolution, a TCP half-close disguised as
2026-08-12 · 12 min read #curation#큐레이션#troubleshooting#postmortem#incident