Resume

Karan Vijayakumar .

Senior Platform Engineer / SRE

Tokyo, Japan Open for opportunities and collaborations
5+
Years Experience
99.99%
Uptime Achieved
10 MW
Accelerator Fleet
2 TB/hr
Telemetry Ingest

— About Me

Professional Summary

Senior SRE / DevOps Engineer with 5+ years building and operating large-scale Kubernetes platforms, CI/CD pipelines, and observability stacks. Specialized in Rust for high-performance infrastructure tooling — custom operators, control loops, and ingestion pipelines handling 100K+ req/s and 2 TB/hr of telemetry. Competitive programmer (Top 700 on LeetCode, top 0.1% globally; 5★ CodeChef) and published researcher in computer vision and deep learning.

— Career

Experience

Mar. 2026 - Present Now

Senior Platform Engineer / SRE

Cogent Labs

  • Built automated reliability tooling for a 10 MW accelerator fleet: DCGM-based health checks surfacing Xid/ECC errors, NVLink faults, and thermal throttling, with a controller that auto-cordons, drains, and recycles degraded nodes before they corrupt inference results or strand capacity
  • Cut GPU-related incident toil and protected goodput by catching silent hardware degradation that standard liveness probes miss
  • Led an end-to-end lights-out GitOps rewrite so every change — from AWS primitive to running pod — is declarative, peer-reviewed, and reconciled automatically with zero manual production access
  • Terraform owns the AWS estate (VPC, IAM/IRSA, EKS, RDS, S3, ElastiCache, MSK) through reusable modules with drift detection; Argo CD drives fleet rollouts with App-of-Apps, ApplicationSets, sync waves, and SLO-gated auto-rollback
  • Eliminated Helm chart sprawl using library charts with values.schema.json validation and shared templates for RBAC, network policies, and OTel sidecars, so new services onboard in hours and on-call no longer touches production
  • Hardened the GitHub Actions supply chain with Cosign image signing, SBOM generation, Grype scanning, and Kyverno admission policies — unsigned or vulnerable artifacts never reach a cluster
Dec. 2024 - Mar. 2026

Senior Site Reliability Engineer

Sales Marker

  • Designed and optimized a hybrid Kubernetes estate (2000+ vCPUs plus specialized GPU nodes) supporting high-throughput backend services, accelerated ML workloads, and data ingestion pipelines, with zero-trust network policies and fine-grained resource partitioning
  • Developed a custom Rust Prometheus-to-Karpenter control loop for workload-aware scheduling, maintaining 98% compute utilization and 99.99% uptime
  • Optimized data shuffle patterns and time-based oversubscription, lifting throughput by 35% and reducing P99 latency by 25% for real-time processing
  • Engineered a high-throughput gRPC API gateway in Rust on Tokio, Tonic, Volo, and Tower handling 100K+ requests/second with sub-millisecond overhead, adding OAuth 2.0 authentication, policy-based rate limiting, and real-time telemetry export
  • Built a self-hosted LGTM stack (Loki/Grafana/Tempo/Mimir) ingesting 2+ TB/hour of telemetry, leading the migration off Datadog on OpenTelemetry standards with custom Rust log aggregators
  • Architected a GitOps operator on kube-rs provisioning ephemeral branch-based sandboxes (RDS/S3/cache) in under 3 minutes, accelerating iteration cycles 4× and catching 3 critical production bugs pre-release
  • Reduced deployment lead time from days to 3 hours through a Grype-scanned GitHub Actions and Argo CD pipeline while holding change-failure rate under 2%
  • Led chaos engineering and quarterly game-days; established SLIs/SLOs and automated guardrails that reduced repeat failures by 37%, consistently meeting a 30-minute RTO and 15-minute RPO
Feb. 2024 - Dec. 2024

Site Reliability Engineer / InfraOps

Rapyuta Robotics

  • Architected and managed cloud-based Kubernetes environments on GCP, implementing GitOps practices with ArgoCD for automated deployments alongside container registry and cloud-native secrets management
  • Engineered custom modules in Python and Go to optimize Kubernetes cluster performance, and implemented service mesh solutions for traffic management, security, and observability
  • Designed a comprehensive observability stack on Prometheus and Grafana with custom exporters, plus anomaly detection via Google Cloud Operations for proactive issue detection
  • Engineered a high-performance Rust eBPF monitoring agent on the Aya library, capturing kernel-level CPU scheduling latency, disk I/O, and socket metrics across a fleet of autonomous robots at under 1% CPU overhead
  • Implemented SLOs/SLIs with custom dashboards and alerting, and optimized data pipelines using Apache Airflow on Kubernetes across cloud data lakes
  • Developed tooling to streamline algorithm production from inception to productization, serving both data scientists and business teams
Jul. 2023 - Feb. 2024

Site Reliability Engineer

FourKites, Inc.

  • Engineered high-performance rate limiting service in Rust handling 100K requests/second with gRPC integration
  • Developed custom layer-4 router in Rust with automated traffic routing, load balancing, and failover logic
  • Built Prometheus exporters in Rust achieving sub-millisecond latency while processing 50K+ metrics/second
  • Implemented LLM-based anomaly detection model in Python trained on 5+ years of telemetry data, improving precision by 15%
  • Designed Go-based microservices architecture for LLM operationalization, processing 95+ time series sources for root cause analysis
  • Optimized data pipelines using Apache Airflow on Kubernetes for scalable analytics across cloud-based data lakes
Aug. 2022 - Jun. 2023

Site Reliability Engineer Intern

FourKites

  • Launched system-wide API protection within 5 months that increased SLA to 5 nines (99.999%)
  • Created plug-and-play rate limiter and improved inter-service call visibility, lowering P1 incidents by over 30%
Nov. 2021 - Apr. 2022

Blockchain Developer Intern

IntelliConnect Technologies

  • Architected end-to-end supply chain management solution using custom Substrate blockchain from scratch
  • Developed core blockchain functionalities including consensus mechanisms, governance, and staking
  • Engineered Smart Contracts using ink! language for product tracking across supply chain
  • Designed integration between IoT devices, ERP systems, and on-chain blockchain via public APIs
  • Deployed Substrate chain on dedicated testnet after security audits and stability testing
Sep. 2021 - Nov. 2021

Software Developer Intern

Fleo.IO

  • Architected full-stack KPI analytics platform with Rust backend processing billions of data points
  • Built Rust-based stream processing framework reducing lag by 60% and React dashboards for real-time visualization
  • Implemented microservices architecture with Docker/Kubernetes, achieving 5 nines uptime with automated failovers
  • Scaled query capacity 10× through intelligent caching and reduced infrastructure costs by 35%
  • Set up CI/CD pipelines using GitLab CI with comprehensive logging and monitoring using Prometheus/Grafana
Apr. 2020 - Jan. 2021

Software Developer Intern

SBNA Technology

  • Formulated ML-based server auto-scaling to improve efficiency, reducing downtime and cost by 5×
  • Streamlined load testing using K6 for scalability, improving SLA by 50%

— Stack

Technical Skills

Infrastructure

09
KubernetesDockerAWSGCPAzureTerraformArgoCDLinuxHelm

Languages

09
RustPythonGolangJavaC#C++TypeScriptJavaScriptBash

Observability

06
PrometheusGrafanaDataDogOpenTelemetryLokiTempo

ML/AI

03
TensorFlowPyTorchScikit-learn

Web & Frontend

04
ReactReact NativeAngularSvelte

Other

18
GitGitHub ActionsGitLab CIPostgreSQLRedisApache AirflowApache SparkHadoopKafkaNode.js.NETGraphQLKarpenterSQLeBPFMimirTableauPower BI

Soft Skills

System DesignArchitectureProblem SolvingLeadershipCommunicationIncident ManagementBlameless PostmortemsCross-functional CollaborationMentorshipTechnical DocumentationPerformance OptimizationCost OptimizationChaos EngineeringDisaster RecoveryScalability Planning

— Showcase

Featured Projects

Unified Image Captioning using Transformers

2021 - 2022

Research Project

  • Designed real-time transformer model for image captioning with custom architecture combining convolutional and recurrent layers (1B+ parameters)
  • Trained novel transformer model on 200,000 images achieving BLEU score of 39

Glaucoma Detection using Vision Transformers

2021 - 2022

Research Project

  • Architected custom real-time vision transformer model for glaucoma detection from fundus images
  • Attained 98.5% accuracy on generalized data after 3 months of training and dedicated R&D

Frames Social Gallery App

2020 - 2021

Team Project

  • Designed social gallery app using React Native, Node.js, .NET, Rust, GraphQL, SQL, and Redis with team of 10
  • Adopted GraphQL and SQL for API design, accelerating deployment time by 70%
  • Launched on iOS and Android with 1,000+ downloads and positive user feedback

— Academic

Education

B.Tech. in Computer Science and Engineering (AI & ML specialization)

Sri Ramachandra Medical College and Research Institute

2019 - 2023

CGPA: 9.33 | Top 5% of university

— Credentials

Certifications

ML Engineering for Production (MLOps)

Specialization

Machine Learning Specialization

Stanford Online

Deep Learning Specialization

DeepLearning.AI

TensorFlow Developer

DeepLearning.AI

Natural Language Processing

Specialization

Machine Learning on Google Cloud

Specialization

Mathematics for Machine Learning

Specialization

Julia Scientific Programming

with Honors

Functional Programming in Scala

Specialization

— Highlights

Achievements

Global Rank #2 in Advent of Code 2023 Day 15 Contest
Winner NeoPat Coding Contest (3× winner: Mar 2022 - Aug 2022)
Hack at SRET Runner-up: Built AI mental health assistant in 10 days
LeetCode Top 700 (Top 0.1% globally)
CodeChef 5⭐ (Top 0.5% globally, Top 0.3% in India)
HackerRank 6 Gold Stars (Top 1% in AI)

— Research

Publications

3rd International Conference on Image Processing and Capsule Networks (2022) - Novel architecture for tuberculosis detection from microscopic sputum smear images

6th International Conference on Trends in Electronics and Informatics (2022) - Evaluation of deep learning framework for glaucoma screening and diagnosis

Let's Connect

Open for opportunities and collaborations