About Me
Karan Vijayakumar.
Senior Platform Engineer
Tokyo, Japan
Building resilient systems that scale. I turn infrastructure chaos into operational excellence.
Impact at a Glance
Accelerator Fleet
10 MW
GPU fleet health and self-healing at Cogent Labs
vCPUs Orchestrated
2000+
Hybrid Kubernetes estate with specialized GPU nodes
Telemetry Ingested
2 TB/hr
Self-hosted LGTM stack, migrated off Datadog
Uptime
99.99%
Held at 98% compute utilization
About Me
I'm a Senior Platform Engineer with a passion for building and scaling infrastructure that just works. Based in Tokyo, Japan, I specialize in cloud-native technologies, Kubernetes orchestration, and ML/AI operations.
My journey started with a deep fascination for competitive programming (Top 400 on LeetCode, 5★ on CodeChef), which gave me a strong foundation in algorithms and problem-solving. This evolved into a career building resilient, scalable systems that handle millions of requests.
I've published research in computer vision and deep learning at international conferences (ICIPCN 2022, ICOEI 2022), combining my academic interests with practical engineering to deliver innovative solutions.
Competitive Programming
- LeetCode Top 700 (0.1% globally)
- CodeChef 5★ (Top 0.5% globally)
- HackerRank 6 Gold Stars (Top 1% AI)
- Advent of Code Global Rank #2 (Day 15, 2023)
What I Do
I focus on four core areas that drive reliability and efficiency in modern infrastructure.
Platform & Cloud Infrastructure
01Hybrid Kubernetes estates, from a 2000+ vCPU fleet to a 10 MW accelerator estate. Lights-out GitOps where Terraform owns the AWS estate and Argo CD reconciles every rollout with SLO-gated auto-rollback.
- Kubernetes
- Terraform
- Argo CD
- AWS/GCP
Rust Systems Tooling
02Custom operators, control loops, and ingestion pipelines in Rust: a Tokio/Tonic gRPC gateway at 100K+ req/s with sub-millisecond overhead, a Prometheus-to-Karpenter scheduler, and an eBPF agent on Aya running under 1% CPU overhead.
- Rust
- Tokio
- kube-rs
- eBPF
Observability
03A self-hosted LGTM stack ingesting 2+ TB/hour, migrated from Datadog on OpenTelemetry standards with custom Rust log aggregators. SLIs, SLOs, and automated guardrails that cut repeat failures by 37%.
- Prometheus
- Grafana
- Loki/Tempo/Mimir
- OpenTelemetry
GPU & ML Operations
04DCGM-based fleet health surfacing Xid/ECC errors, NVLink faults, and thermal throttling, with automated node recycling. Earlier, an LLM-based anomaly detection model trained on 5+ years of telemetry that improved precision by 15%.
- DCGM
- PyTorch
- TensorFlow
- MLOps
Technical Skills
A comprehensive toolkit built over years of hands-on experience.
Infrastructure & Cloud
- Kubernetes 95%
- Terraform 90%
- AWS 90%
- Docker 92%
- Linux 92%
Programming
- Rust 95%
- Python 92%
- Go 85%
- TypeScript 85%
- Bash 90%
Observability
- Prometheus 92%
- Grafana 92%
- Loki/Tempo/Mimir 88%
- OpenTelemetry 85%
- eBPF 80%
ML & Data
- PyTorch 85%
- TensorFlow 82%
- Scikit-learn 80%
- Apache Airflow 82%
- Apache Spark 78%
Career Journey
From competitive programming enthusiast to building infrastructure at scale.
2026
PresentSenior Platform Engineer / SRE
Cogent Labs, Tokyo
Reliability tooling for a 10 MW accelerator fleet, and a lights-out GitOps rewrite that removed manual production access entirely.
- 10 MW fleet self-healing
- Lights-out GitOps rewrite
- Cosign/SBOM/Kyverno supply chain
2024
Senior Site Reliability Engineer
Sales Marker, Tokyo
A 2000+ vCPU hybrid Kubernetes estate for high-throughput backends, accelerated ML workloads, and ingestion pipelines.
- 99.99% uptime at 98% utilization
- 100K+ req/s Rust gRPC gateway
- 2 TB/hr LGTM pipeline
2024
Site Reliability Engineer / InfraOps
Rapyuta Robotics, Tokyo
GKE platform engineering with ArgoCD GitOps, service mesh, and kernel-level observability across a robot fleet.
- GKE + ArgoCD GitOps
- Rust eBPF agent under 1% overhead
- Prometheus/Grafana stack
2022
Site Reliability Engineer
FourKites, Chennai
Joined as an SRE intern and moved into the SRE role. Rust infrastructure services and LLM-based anomaly detection.
- 100K req/s Rust rate limiter
- 50K+ metrics/s Prometheus exporters
- Raised SLA to five nines
2020
Software & Blockchain Internships
IntelliConnect, Fleo.IO, SBNA Technology
Early systems work in Rust and Substrate: a custom blockchain for supply chain traceability, and a KPI analytics platform.
- Substrate chain with ink! contracts
- Rust KPI platform, 10× query capacity
- ML-based server autoscaling
2019 — 2023
B.Tech Computer Science (AI & ML)
Sri Ramachandra Medical College and Research Institute
Specialized in AI/ML with published research in computer vision and deep learning.
- CGPA 9.33, top 5% of university
- 2 international conference papers
- Top 0.1% competitive programming
My Engineering Philosophy
The principles that guide how I approach infrastructure and reliability.
- 01
Reliability First
Every system I build prioritizes uptime and resilience. Failures are opportunities for improvement, not blame.
- 02
Data-Driven Decisions
Metrics guide everything. If we can't measure it, we can't improve it. SLOs and error budgets are sacred.
- 03
Collaborative Problem Solving
The best solutions come from diverse perspectives. Blameless postmortems and knowledge sharing are key.
- 04
Automate Everything
Toil is the enemy. If it can be automated, it should be. That frees the team to focus on real problems.
Research & Publications
ICIPCN 2022
Novel architecture for tuberculosis detection from microscopic sputum smear images
ICOEI 2022
Evaluation of deep learning framework for glaucoma screening and diagnosis
Let's Build Something Amazing
Always interested in challenging infrastructure problems and opportunities to make systems more reliable.