Building dependable systems

I make complex
infrastructure behave.

I’m Kaifeng Lin, a software engineer focused on backend systems, Kubernetes, and infrastructure reliability.

01 / Selected work

Systems built for the messy parts.

A few areas I’ve worked deeply in. Short versions for now; longer engineering notes will follow.

02

GPU · Reliability

GPU Fleet Recovery

Automated fault detection, isolation, remediation, and burn-in validation—turning hardware failures into repeatable workflows.

  • NVML
  • Linux
  • Automation
03

Scheduling · Linux

Resource-Aware Scheduling

Making disk I/O visible to scheduling and cgroup enforcement to reduce noisy-neighbor failures in shared clusters.

  • cgroups
  • Scheduler
  • Observability
04

Open Source · Kubernetes

Kubelet Storage Accounting

Reported and root-caused an ephemeral-storage accounting bug affecting multi-container workloads.

View issue

02 / Field notes

Writing from the workbench.

Short, practical notes on infrastructure internals and the lessons that only show up after systems meet production.

First notes are in progress.

I’ll publish them here as they’re ready.

  • 01Kubernetes internals
  • 02Production debugging
  • 03Infrastructure design

03 / About

Engineering with curiosity, evidence, and a healthy respect for failure modes.

I enjoy going below the abstraction: reading source, profiling behavior, and turning one-off incidents into durable systems.