Blog
Writing on GPUs, LLMs, MLOps, Kubernetes — and mindset · 3525 posts
#2026-03 765#english 592#culture 264#deep-dive 254#kubernetes 249#career 229#ai 219#llm 209#devops 196#2026-04 146#security 143#database 114#observability 113#communication 109#history 107#architecture 100#productivity 96#finance 88#economy 84#mindset 81#psychology 80#ai-papers 79#food 78#it 78#travel 78#deep-learning 77#japanese 77#networking 77#performance 72#business-travel 70#linux 70#gpu 69#ai-agent 66#cs-fundamentals 63#postgresql 61#rag 58#self-improvement 55#learning 53#mlops 53#python 52
SOTA Music and Audio Generation — Neural Codecs and Generative Models
A lineage-focused overview from audio representations (waveform, spectrogram, neural codec) to autoregressive audio language models, diffusion-based audio, and text-to-music conditioning. We analyze the principles of the
2026-06-30 · 9 min read #ai-papers#audio-generation#music-generation#neural-codec#audio-language-modelSOTA Video Generation Models Explained — Spatiotemporal Diffusion Transformers
A look at the two fundamental challenges of video generation, temporal consistency and compute cost, through the lens of spatiotemporal latent patches and diffusion transformers. We analyze the concepts Sora introduced,
2026-06-30 · 8 min read #ai-papers#video-generation#diffusion-transformer#spatiotemporal#text-to-videoAnalyzing SOTA Image Generation Models — From Diffusion to FLUX
A lineage-centered overview of the frontier of text-to-image generation, from diffusion model fundamentals through latent diffusion, DiT, rectified flow, and the FLUX family. We analyze the shared structure and differenc
2026-06-30 · 10 min read #ai-papers#diffusion-models#text-to-image#latent-diffusion#rectified-flowSOTA Real-Time Video Analysis — Tracking, Understanding, Efficient Inference ♪ Listenable
A survey of SOTA trends in real-time video analysis. We cover video understanding tasks like action recognition, object tracking, and temporal segmentation, the SAM 2 family concept of video segmentation and tracking, tr
2026-06-30 · 21 min read #ai-papers#video-understanding#object-tracking#sam2#action-recognitionRust Smart Pointers: Box, Rc, Arc, RefCell & Interior Mutability
Smart pointers unlock the cases Rust ownership alone struggles to express. Put values on the heap with Box, share ownership with Rc and its thread-safe cousin Arc, move the borrow rules from compile time to run time with
2026-06-30 · 11 min read #rust#memory#systemsNaming Things: The Hardest Easy Skill in Code
The joke that "there are only two hard things in computer science: cache invalidation and naming things" is only half a joke. Why a good name reveals intent rather than implementation, how name length scales with scope,
2026-06-30 · 8 min read #clean-code#engineering#fundamentalsThe FDE Playbook: Winning With Throwaway Prototypes
How a forward deployed engineer flips the table by shipping a working demo on the customer real data in days. Deliberate, disposable tech debt; the spike that de-risks; when to hardcode; getting to the aha moment fast; m
2026-06-30 · 8 min read #engineering#prototyping#fdeDatabase Transactions and Isolation Levels, Explained
What ACID actually guarantees, the four isolation levels (read uncommitted → serializable), the anomalies each one permits (dirty, non-repeatable, and phantom reads, plus write skew), MVCC, locking versus optimistic conc
2026-06-30 · 13 min read #databases#transactions#acidAnalyzing SOTA Multimodal LLMs — One Model to See, Hear, and Speak ♪ Listenable
How did a language model trained purely on text come to understand and generate images, audio, and video? This post walks through modality encoders and projectors, the unified token space, the any-to-any flow, native mul
2026-06-30 · 22 min read #multimodal-llm#any-to-any#vision-language#audio#architectureSOTA Speech Recognition and Synthesis — From Whisper to Codec Language Models
We survey recent trends in speech recognition (ASR) and synthesis (TTS). From HMM to CTC/attention, Whisper large-scale weak supervision, and from Tacotron to neural vocoders and codec language models, we trace the linea
2026-06-30 · 15 min read #ai-papers#speech-recognition#text-to-speech#whisper#neural-codecSOTA Autonomous Driving Perception — BEV, Occupancy, End-to-End
We survey recent trends in autonomous driving perception. From the perception-prediction-planning-control stack, BEV multi-camera fusion, occupancy networks 3D occupancy representation, vision-centric vs LiDAR fusion, to
2026-06-30 · 16 min read #ai-papers#autonomous-driving#bev#occupancy-network#end-to-endRobot Safety and Alignment — Trusting Powerful Robots
How can we trust increasingly powerful robots. We take a balanced look at physical safety, constrained reinforcement learning and safety layers, handling distribution shift, human-robot collaboration safety, verification
2026-06-29 · 17 min read #ai-papers#robotics#safety#alignment#reinforcement-learningRobots That Learn from Human Video — The Dream of Web-Scale Data ♪ Listenable
Can a robot learn from videos of people handling objects. We cover affordances and trajectories, the domain gap, representation learning and pre-training, one-shot imitation, web-video scale-up, and combining with robot
2026-06-29 · 19 min read #ai-papers#robotics#imitation-learning#representation-learning#videoRobot Foundation Models — One Policy for Many Jobs ♪ Listenable
We lay out the trend of robot foundation models that aim to do many jobs across many robots with a single policy. We cover the concept of a generalist policy, large-scale robot data such as Open X-Embodiment, cross-embod
2026-06-29 · 19 min read #ai-papers#robotics#foundation-model#generalist-policy#open-x-embodimentWorld Models — When Robots Imagine the Future
An explanation of the concept of world models (learning environment dynamics), covering model-based reinforcement learning, latent-space prediction, the robotic application of video prediction and generative models, plan
2026-06-29 · 17 min read #ai-papers#robotics#world-models#model-based-rl#planningSim-to-Real — Bringing What Was Learned in Simulation into Reality
An analysis of why robot policies learned in simulation collapse in reality (the reality gap), covering countermeasures such as domain randomization, domain adaptation, system identification, and digital twins, along wit
2026-06-29 · 17 min read #ai-papers#robotics#sim-to-real#domain-randomization#digital-twinHow Robots Learn — Imitation Learning and Reinforcement Learning
An overview of the four ways robots acquire skills, followed by a deep look at imitation learning (teleoperation, behavioral cloning, DAgger) and reinforcement learning (rewards, policies, exploration): their principles,
2026-06-29 · 16 min read #ai-papers#robotics#imitation-learning#reinforcement-learning#robot-learningThe Robots Eye — 3D Perception and SLAM ♪ Listenable
How robots see the world. We walk through RGB-D, LiDAR, and stereo sensors, the SLAM pipeline, point clouds and voxels, 6D pose estimation, deep-learning perception, and finally NeRF and 3D Gaussian Splatting, with code
2026-06-29 · 18 min read #ai-papers#robotics#slam#3d-perception#pointcloudHumanoid Whole-Body Control — Walking on Two Legs and Handling with Two Hands ♪ Listenable
From walking on two legs to handling objects with hands, we lay out the core ideas of bipedal locomotion and whole-body control. We cover ZMP and MPC, reinforcement-learning locomotion, balance and fall recovery, the int
2026-06-29 · 21 min read #ai-papers#robotics#humanoid#locomotion#whole-body-controlRegex From Zero to Confident: Character Classes, Quantifiers, Anchors, Groups, Lookarounds, and ReDoS
If regular expressions have always looked like line noise, let this post fix that. We build up the pieces one at a time — character classes, quantifiers, anchors, groups, alternation, lookarounds — then cover greedy vers
2026-06-29 · 10 min read #regex#programming#fundamentals