Resume
Karan Vijayakumar .
Senior Platform Engineer / SRE
— About Me
Professional Summary
Senior SRE / DevOps Engineer with 5+ years building and operating large-scale Kubernetes platforms, CI/CD pipelines, and observability stacks. Specialized in Rust for high-performance infrastructure tooling — custom operators, control loops, and ingestion pipelines handling 100K+ req/s and 2 TB/hr of telemetry. Competitive programmer (Top 700 on LeetCode, top 0.1% globally; 5★ CodeChef) and published researcher in computer vision and deep learning.
— Career
Experience
Senior Platform Engineer / SRE
Cogent Labs
- Built automated reliability tooling for a 10 MW accelerator fleet: DCGM-based health checks surfacing Xid/ECC errors, NVLink faults, and thermal throttling, with a controller that auto-cordons, drains, and recycles degraded nodes before they corrupt inference results or strand capacity
- Cut GPU-related incident toil and protected goodput by catching silent hardware degradation that standard liveness probes miss
- Led an end-to-end lights-out GitOps rewrite so every change — from AWS primitive to running pod — is declarative, peer-reviewed, and reconciled automatically with zero manual production access
- Terraform owns the AWS estate (VPC, IAM/IRSA, EKS, RDS, S3, ElastiCache, MSK) through reusable modules with drift detection; Argo CD drives fleet rollouts with App-of-Apps, ApplicationSets, sync waves, and SLO-gated auto-rollback
- Eliminated Helm chart sprawl using library charts with values.schema.json validation and shared templates for RBAC, network policies, and OTel sidecars, so new services onboard in hours and on-call no longer touches production
- Hardened the GitHub Actions supply chain with Cosign image signing, SBOM generation, Grype scanning, and Kyverno admission policies — unsigned or vulnerable artifacts never reach a cluster
Senior Site Reliability Engineer
Sales Marker
- Designed and optimized a hybrid Kubernetes estate (2000+ vCPUs plus specialized GPU nodes) supporting high-throughput backend services, accelerated ML workloads, and data ingestion pipelines, with zero-trust network policies and fine-grained resource partitioning
- Developed a custom Rust Prometheus-to-Karpenter control loop for workload-aware scheduling, maintaining 98% compute utilization and 99.99% uptime
- Optimized data shuffle patterns and time-based oversubscription, lifting throughput by 35% and reducing P99 latency by 25% for real-time processing
- Engineered a high-throughput gRPC API gateway in Rust on Tokio, Tonic, Volo, and Tower handling 100K+ requests/second with sub-millisecond overhead, adding OAuth 2.0 authentication, policy-based rate limiting, and real-time telemetry export
- Built a self-hosted LGTM stack (Loki/Grafana/Tempo/Mimir) ingesting 2+ TB/hour of telemetry, leading the migration off Datadog on OpenTelemetry standards with custom Rust log aggregators
- Architected a GitOps operator on kube-rs provisioning ephemeral branch-based sandboxes (RDS/S3/cache) in under 3 minutes, accelerating iteration cycles 4× and catching 3 critical production bugs pre-release
- Reduced deployment lead time from days to 3 hours through a Grype-scanned GitHub Actions and Argo CD pipeline while holding change-failure rate under 2%
- Led chaos engineering and quarterly game-days; established SLIs/SLOs and automated guardrails that reduced repeat failures by 37%, consistently meeting a 30-minute RTO and 15-minute RPO
Site Reliability Engineer / InfraOps
Rapyuta Robotics
- Architected and managed cloud-based Kubernetes environments on GCP, implementing GitOps practices with ArgoCD for automated deployments alongside container registry and cloud-native secrets management
- Engineered custom modules in Python and Go to optimize Kubernetes cluster performance, and implemented service mesh solutions for traffic management, security, and observability
- Designed a comprehensive observability stack on Prometheus and Grafana with custom exporters, plus anomaly detection via Google Cloud Operations for proactive issue detection
- Engineered a high-performance Rust eBPF monitoring agent on the Aya library, capturing kernel-level CPU scheduling latency, disk I/O, and socket metrics across a fleet of autonomous robots at under 1% CPU overhead
- Implemented SLOs/SLIs with custom dashboards and alerting, and optimized data pipelines using Apache Airflow on Kubernetes across cloud data lakes
- Developed tooling to streamline algorithm production from inception to productization, serving both data scientists and business teams
Site Reliability Engineer
FourKites, Inc.
- Engineered high-performance rate limiting service in Rust handling 100K requests/second with gRPC integration
- Developed custom layer-4 router in Rust with automated traffic routing, load balancing, and failover logic
- Built Prometheus exporters in Rust achieving sub-millisecond latency while processing 50K+ metrics/second
- Implemented LLM-based anomaly detection model in Python trained on 5+ years of telemetry data, improving precision by 15%
- Designed Go-based microservices architecture for LLM operationalization, processing 95+ time series sources for root cause analysis
- Optimized data pipelines using Apache Airflow on Kubernetes for scalable analytics across cloud-based data lakes
Site Reliability Engineer Intern
FourKites
- Launched system-wide API protection within 5 months that increased SLA to 5 nines (99.999%)
- Created plug-and-play rate limiter and improved inter-service call visibility, lowering P1 incidents by over 30%
Blockchain Developer Intern
IntelliConnect Technologies
- Architected end-to-end supply chain management solution using custom Substrate blockchain from scratch
- Developed core blockchain functionalities including consensus mechanisms, governance, and staking
- Engineered Smart Contracts using ink! language for product tracking across supply chain
- Designed integration between IoT devices, ERP systems, and on-chain blockchain via public APIs
- Deployed Substrate chain on dedicated testnet after security audits and stability testing
Software Developer Intern
Fleo.IO
- Architected full-stack KPI analytics platform with Rust backend processing billions of data points
- Built Rust-based stream processing framework reducing lag by 60% and React dashboards for real-time visualization
- Implemented microservices architecture with Docker/Kubernetes, achieving 5 nines uptime with automated failovers
- Scaled query capacity 10× through intelligent caching and reduced infrastructure costs by 35%
- Set up CI/CD pipelines using GitLab CI with comprehensive logging and monitoring using Prometheus/Grafana
Software Developer Intern
SBNA Technology
- Formulated ML-based server auto-scaling to improve efficiency, reducing downtime and cost by 5×
- Streamlined load testing using K6 for scalability, improving SLA by 50%
— Stack
Technical Skills
Infrastructure
Languages
Observability
ML/AI
Web & Frontend
Other
Soft Skills
— Showcase
Featured Projects
Unified Image Captioning using Transformers
2021 - 2022Research Project
- Designed real-time transformer model for image captioning with custom architecture combining convolutional and recurrent layers (1B+ parameters)
- Trained novel transformer model on 200,000 images achieving BLEU score of 39
Glaucoma Detection using Vision Transformers
2021 - 2022Research Project
- Architected custom real-time vision transformer model for glaucoma detection from fundus images
- Attained 98.5% accuracy on generalized data after 3 months of training and dedicated R&D
Frames Social Gallery App
2020 - 2021Team Project
- Designed social gallery app using React Native, Node.js, .NET, Rust, GraphQL, SQL, and Redis with team of 10
- Adopted GraphQL and SQL for API design, accelerating deployment time by 70%
- Launched on iOS and Android with 1,000+ downloads and positive user feedback
— Academic
Education
B.Tech. in Computer Science and Engineering (AI & ML specialization)
Sri Ramachandra Medical College and Research Institute
CGPA: 9.33 | Top 5% of university
— Credentials
Certifications
ML Engineering for Production (MLOps)
SpecializationMachine Learning Specialization
Stanford OnlineDeep Learning Specialization
DeepLearning.AITensorFlow Developer
DeepLearning.AINatural Language Processing
SpecializationMachine Learning on Google Cloud
SpecializationMathematics for Machine Learning
SpecializationJulia Scientific Programming
with HonorsFunctional Programming in Scala
Specialization— Highlights
Achievements
— Research
Publications
3rd International Conference on Image Processing and Capsule Networks (2022) - Novel architecture for tuberculosis detection from microscopic sputum smear images
6th International Conference on Trends in Electronics and Informatics (2022) - Evaluation of deep learning framework for glaucoma screening and diagnosis