Karan Vijayakumar Available

About Me

Karan Vijayakumar.

Senior Platform Engineer

Tokyo, Japan

Building resilient systems that scale. I turn infrastructure chaos into operational excellence.

Impact at a Glance

Accelerator Fleet

10 MW

GPU fleet health and self-healing at Cogent Labs

vCPUs Orchestrated

2000+

Hybrid Kubernetes estate with specialized GPU nodes

Telemetry Ingested

2 TB/hr

Self-hosted LGTM stack, migrated off Datadog

Uptime

99.99%

Held at 98% compute utilization

About Me

I'm a Senior Platform Engineer with a passion for building and scaling infrastructure that just works. Based in Tokyo, Japan, I specialize in cloud-native technologies, Kubernetes orchestration, and ML/AI operations.

My journey started with a deep fascination for competitive programming (Top 400 on LeetCode, 5★ on CodeChef), which gave me a strong foundation in algorithms and problem-solving. This evolved into a career building resilient, scalable systems that handle millions of requests.

I've published research in computer vision and deep learning at international conferences (ICIPCN 2022, ICOEI 2022), combining my academic interests with practical engineering to deliver innovative solutions.

Competitive Programming

  • LeetCode Top 700 (0.1% globally)
  • CodeChef 5★ (Top 0.5% globally)
  • HackerRank 6 Gold Stars (Top 1% AI)
  • Advent of Code Global Rank #2 (Day 15, 2023)

What I Do

I focus on four core areas that drive reliability and efficiency in modern infrastructure.

Platform & Cloud Infrastructure

01

Hybrid Kubernetes estates, from a 2000+ vCPU fleet to a 10 MW accelerator estate. Lights-out GitOps where Terraform owns the AWS estate and Argo CD reconciles every rollout with SLO-gated auto-rollback.

  • Kubernetes
  • Terraform
  • Argo CD
  • AWS/GCP

Rust Systems Tooling

02

Custom operators, control loops, and ingestion pipelines in Rust: a Tokio/Tonic gRPC gateway at 100K+ req/s with sub-millisecond overhead, a Prometheus-to-Karpenter scheduler, and an eBPF agent on Aya running under 1% CPU overhead.

  • Rust
  • Tokio
  • kube-rs
  • eBPF

Observability

03

A self-hosted LGTM stack ingesting 2+ TB/hour, migrated from Datadog on OpenTelemetry standards with custom Rust log aggregators. SLIs, SLOs, and automated guardrails that cut repeat failures by 37%.

  • Prometheus
  • Grafana
  • Loki/Tempo/Mimir
  • OpenTelemetry

GPU & ML Operations

04

DCGM-based fleet health surfacing Xid/ECC errors, NVLink faults, and thermal throttling, with automated node recycling. Earlier, an LLM-based anomaly detection model trained on 5+ years of telemetry that improved precision by 15%.

  • DCGM
  • PyTorch
  • TensorFlow
  • MLOps

Technical Skills

A comprehensive toolkit built over years of hands-on experience.

01

Infrastructure & Cloud

  • Kubernetes 95%
  • Terraform 90%
  • AWS 90%
  • Docker 92%
  • Linux 92%
02

Programming

  • Rust 95%
  • Python 92%
  • Go 85%
  • TypeScript 85%
  • Bash 90%
03

Observability

  • Prometheus 92%
  • Grafana 92%
  • Loki/Tempo/Mimir 88%
  • OpenTelemetry 85%
  • eBPF 80%
04

ML & Data

  • PyTorch 85%
  • TensorFlow 82%
  • Scikit-learn 80%
  • Apache Airflow 82%
  • Apache Spark 78%

Career Journey

From competitive programming enthusiast to building infrastructure at scale.

  1. 2026

    Present

    Senior Platform Engineer / SRE

    Cogent Labs, Tokyo

    Reliability tooling for a 10 MW accelerator fleet, and a lights-out GitOps rewrite that removed manual production access entirely.

    • 10 MW fleet self-healing
    • Lights-out GitOps rewrite
    • Cosign/SBOM/Kyverno supply chain
  2. 2024

    Senior Site Reliability Engineer

    Sales Marker, Tokyo

    A 2000+ vCPU hybrid Kubernetes estate for high-throughput backends, accelerated ML workloads, and ingestion pipelines.

    • 99.99% uptime at 98% utilization
    • 100K+ req/s Rust gRPC gateway
    • 2 TB/hr LGTM pipeline
  3. 2024

    Site Reliability Engineer / InfraOps

    Rapyuta Robotics, Tokyo

    GKE platform engineering with ArgoCD GitOps, service mesh, and kernel-level observability across a robot fleet.

    • GKE + ArgoCD GitOps
    • Rust eBPF agent under 1% overhead
    • Prometheus/Grafana stack
  4. 2022

    Site Reliability Engineer

    FourKites, Chennai

    Joined as an SRE intern and moved into the SRE role. Rust infrastructure services and LLM-based anomaly detection.

    • 100K req/s Rust rate limiter
    • 50K+ metrics/s Prometheus exporters
    • Raised SLA to five nines
  5. 2020

    Software & Blockchain Internships

    IntelliConnect, Fleo.IO, SBNA Technology

    Early systems work in Rust and Substrate: a custom blockchain for supply chain traceability, and a KPI analytics platform.

    • Substrate chain with ink! contracts
    • Rust KPI platform, 10× query capacity
    • ML-based server autoscaling
  6. 2019 — 2023

    B.Tech Computer Science (AI & ML)

    Sri Ramachandra Medical College and Research Institute

    Specialized in AI/ML with published research in computer vision and deep learning.

    • CGPA 9.33, top 5% of university
    • 2 international conference papers
    • Top 0.1% competitive programming

My Engineering Philosophy

The principles that guide how I approach infrastructure and reliability.

  • 01

    Reliability First

    Every system I build prioritizes uptime and resilience. Failures are opportunities for improvement, not blame.

  • 02

    Data-Driven Decisions

    Metrics guide everything. If we can't measure it, we can't improve it. SLOs and error budgets are sacred.

  • 03

    Collaborative Problem Solving

    The best solutions come from diverse perspectives. Blameless postmortems and knowledge sharing are key.

  • 04

    Automate Everything

    Toil is the enemy. If it can be automated, it should be. That frees the team to focus on real problems.

Research & Publications

  • ICIPCN 2022

    Novel architecture for tuberculosis detection from microscopic sputum smear images

  • ICOEI 2022

    Evaluation of deep learning framework for glaucoma screening and diagnosis

Let's Build Something Amazing

Always interested in challenging infrastructure problems and opportunities to make systems more reliable.