CV
Experience
Kubernetes Developer @ Bilibili, Container Team
2025 – present
- Multi‑cluster scheduling (KubeAir) – Optimized
kubeair-agentperformance to ensure stable API response even as vpool count grows; added support for specifying GPU card types in fixed resource pools (previously only one type per pool); participated in vpool tag governance and resource reporting bug fixes. - Karmada integration & SLO monitoring – Developed workload SLO health checks for
deployment/karmada-deployment, monitoring control plane availability for both clusters and Karmada layer; contributed to multi‑cluster scheduling improvements. - Cluster monitoring component (k8s-cluster-monitor) – Designed and implemented from scratch: monitors node/pod counts with alerts, schedulable instance inspection (unipool 8c8g in shwgq/shylf/shjd) with WeCom notifications, and certificate expiration checks. Bug fixes for hub‑checker health check methods.
- AI device support (NPU/GPU) – Developed
npu-exporterfor Huawei Ascend NPU container‑level metrics, adapted to company monitoring platform; managedgpu-manager/gpu-monitoringreleases, fixed GPU resource reporting bugs, standardized multi‑cluster YAML for gpu-manager; extended gpu-monitor to support SM‑level and container‑level GPU utilization metrics. - Autoscaling (VPA/HPA) – Maintained and developed VPA/HPA components; introduced quota‑aware VPA (ensuring desired values stay within quota); fixed multiple bugs; added quota restrictions to VPA and k8s-webhook for soft/hard limit governance.
- Operational support – Wrote diff scripts for unipool capacity data across months; handled machine replacement/return/migration (CNY resource eviction, 204/402 data center relocation); new cluster setup; scheduling side on-call; monitoring governance for k8s master nodes and Karmada control plane components.
Intern @ NetEase – Computational Intelligence Group
Jun 2024 – Dec 2024
- Assisted in GPU hot‑plug feature development for Kubernetes clusters.
- Developed a message queue mechanism for scheduler callbacks, improving scheduling event handling.
- Tech stack: Kubernetes, Docker, Go.
Key Contributions & Quantified Impact
| Focus Area | Contribution | Weight |
|---|---|---|
| Multi‑cluster scheduling (KubeAir) | Performance optimization, card type differentiation, vpool governance | 30% |
| Cluster monitoring component | Independent design & rollout – workload/capacity/certificate health checks | 30% |
| AI device components (NPU/GPU) | npu-exporter, gpu-manager/gpu-monitoring maintenance & iteration | 30% |
| Autoscaling (VPA/HPA) | Quota integration, bug fixes, ensuring reliable scaling within limits | 10% |
Open Source & Community
- Aiming to become an active Karmada contributor – currently exploring multi‑cluster scheduling and federation, with hands‑on experience in Karmada SLO monitoring and deployment health checks.
- Interested in contributing to Karmada’s scheduling policies and multi‑cluster autoscaling.
Education
- M.S. in Computer Science, Zhejiang University (2022–2025)
Advised by Dr. Kai Bu - B.S. in Computer Science, Hangzhou Dianzi University (2018–2022)
Skills
- Languages: Go, Python, Java
- Cloud & Orchestration: Kubernetes, Karmada, Docker, Multi‑cluster Scheduling (KubeAir)
- Monitoring & Autoscaling: Prometheus, HPA, VPA, GPU/NPU metrics
- Tools: Git, Linux, CI/CD, WeCom API
Projects (Selected)
More details available upon request – contact via GitHub or email.
