Andrei Solodov

Platform engineer focused on reliability, automation, and observability.

About

My major is cryptography. My diploma states that i'm a mathematician. I can play Fur Elise on piano (first movement plus a bit more). Certified Kubernetes Administrator. I'm mostly interested to find the job (if i'm lucky but i'm not in a hurry) in adtech or blockchain, fintech startups.

Location: Remote, US

Education: Math

Cert: Certified Kubernetes Administrator

Hobby: Piano, Skateboarding

Experience

Senior Platform Engineer — AppLovin

Nov 2022 - Apr 2026 · Palo Alto, California

  • - Managing and supporting hundreds high-load Kubernetes clusters (many thousands nodes) on GCP GKE, ensuring optimal performance and reliability for real-time systems (bidding, mediation, inference)
  • - Setting up continuous profiling solutions: Grafana Pyroscope, Yandex Perforator
  • - Setting up mission critical massive monitoring system on VictoriaMetrics (hundreds of billions active time-series)
  • - Rollout massive GPU based Kubernetes clusters to run inference workloads (NVIDIA L4, V100)
  • - Rollout training compute instances with NVIDIA H100/200 to support research science teams to work with PyTorch, Tensorflow, cuDNN
  • - Supported deployment of multiple ML models using Docker and Kubernetes
  • - Integrated llama.cpp for on-device language model inference in a production environment
  • - Maintained CI/CD pipelines for model releases
  • - Monitored model performance and errors through logs and dashboards
  • - Worked with Weights & Biases and Hugging Face in an ML team environment
  • - Maintaining robust observability systems utilizing VictoriaMetrics, Prometheus, Loki, and Quickwit, enabling efficient monitoring and troubleshooting.
  • - Implementing and optimizing CI/CD pipelines using tools like Jenkins, GitHub Actions, and Bazel (huge C++ applications), improving build and deployment workflows.
  • - Automating tasks and building internal system utilities in Go, enhancing operational efficiency and reducing manual interventions.
  • - Administering Aerospike clusters, ensuring high availability and performance for distributed database workloads.
  • - Managing deployments using ArgoCD, streamlining GitOps workflows and maintaining consistency across Kubernetes environments.
  • - Leveraging Google Cloud Platform (GCP) services and infrastructure to support scalable and cost-efficient operations.
  • - Actively working on open source projects. Some examples: https://github.com/yandex/perforator/pull/36 , https://github.com/VictoriaMetrics/operator/pull/763

Staff Software Engineer — EngineYard

Nov 2021 - Apr 2022 · Texas

  • Continuous development and improvement of a NoOps Engine Yard Kontainers PaaS based on AWS EKS.
  • Designing internal platform services, CLI and IaaC. Customer features development as part of a Professional Services offering. MWs preparation and execution. RCA, deep-dives preparation. L3 on-call in case of a major issue

Senior Site Reliability Engineer — Aurea

Feb 2016 - Nov 2021 · Texas

  • Designed and supported internal platforms: Kubernetes v1.22 (runc + containerd on Amazon Linux 2 spot nodes, AMD64 and ARM, with cluster autoscaler, HPA, VPA) and Consul clusters, plus 10+ vanilla Docker clusters on AWS x1-family EC2 (including Windows Containers)
  • Handled incident and change management;
  • Prepared post-mortems/RCAs using strace, Docker Go crash stack traces, pprof, and Mark Russinovich's Sysinternals toolset for Windows Servers
  • Supported 300+ company AWS accounts across multiple AWS Orgs with SCPs
  • Developed an internal service stack: CLIs (AWS SDK for Go) and JSON/HTTP applications (mainly Go)
  • Designed CI/CD pipelines (Jenkins, TeamCity) and custom monitoring systems (Kafka, Prometheus/Thanos, custom Grafana dashboards) integrated with PagerDuty and Atlassian Opsgenie
  • Served as L3 on-call, the last resort during incidents
  • Implemented container security: image vulnerability scanning, LDAP-based RBAC
  • Designed a Terraform + Ansible IaC framework enabling rapid system deployment as part of disaster recovery planning
  • Supported lift-and-shift migrations from on-prem datacenters to AWS and Kubernetes
  • Applied advanced networking expertise: AWS VPC, Direct Connect, S2S VPN, BGP, NAT
  • Configured software routers (VyOS, Openswan, PaloAlto) to connect public clouds, often across overlapping CIDRs
  • Built 50+ TB storage layers with varying IO/latency needs: LVM + ext4/xfs, ZFS, NFS/NetApp

Skills

Orchestration

  • Kubernetes
  • Docker
  • Helm
  • Argo CD

Cloud & Infra

  • AWS
  • EKS
  • Terraform
  • Linux
  • Networking

Education

Samara State University — Math (2003 - 2009)