Kubernetes · Scalability
Control Plane at Scale
Profiling API-server and storage bottlenecks, shaping request flows, and building SLO-driven guardrails for large clusters.
Building dependable systems
I’m Kaifeng Lin, a software engineer focused on backend systems, Kubernetes, and infrastructure reliability.
01 / Selected work
A few areas I’ve worked deeply in. Short versions for now; longer engineering notes will follow.
Kubernetes · Scalability
Profiling API-server and storage bottlenecks, shaping request flows, and building SLO-driven guardrails for large clusters.
GPU · Reliability
Automated fault detection, isolation, remediation, and burn-in validation—turning hardware failures into repeatable workflows.
Scheduling · Linux
Making disk I/O visible to scheduling and cgroup enforcement to reduce noisy-neighbor failures in shared clusters.
Open Source · Kubernetes
Reported and root-caused an ephemeral-storage accounting bug affecting multi-container workloads.
View issue02 / Field notes
Short, practical notes on infrastructure internals and the lessons that only show up after systems meet production.
I’ll publish them here as they’re ready.
03 / About
I enjoy going below the abstraction: reading source, profiling behavior, and turning one-off incidents into durable systems.