Bengaluru, India

Suraj C Jagatap

Site Reliability Engineer

I build cloud infrastructure and keep it running when it would rather not.

p95 response time

About

For seven years I've built and looked after infrastructure on AWS — provisioning it, automating it, watching it, and fixing it when something breaks at three in the morning.

Most of that work has been in distribution and retail, where a few hours of downtime means trucks don't move and orders don't ship. That taught me to care about recovery more than elegance. A system that fails quietly and comes back fast beats a clever one that surprises you.

I'm self-taught. Everything I know came from building things, breaking them, and being on call for the result. These days my time goes on multi-account AWS setups, Kubernetes, Terraform, and cutting cloud bills that grew while nobody was watching.

0years on call
0AWS accounts
ap-south-1home region

Experience

  • Now

    Senior Software Engineer, Site Reliability

    Infinite Locus

    Running infrastructure across several client accounts — provisioning, monitoring, cost work, and the day-to-day of keeping other people's production systems healthy.

    AWSKubernetesTerraform
  • Previously

    Infrastructure & Hosting

    Intelligent Retail (Ripplr)

    Owned hosting and infrastructure for an FMCG distribution platform — the systems that move stock and orders for a physical supply chain.

    AWSDockerLinux
  • Earlier

    Operations & Logistics Systems

    Aditya Birla

    Inventory and inward/outward flows, internal file transfer, and keeping cloud APIs healthy. Where I learned what downtime costs a business that ships real things.

    OperationsAPIs

Things I built

  • Internal

    Multi-account health monitor

    One dashboard showing the health of every AWS account we manage, instead of logging into each one to find out something had already broken.

    PythonCloudWatchboto3
  • Mobile

    PhysioBook

    A scheduling and records app for a home-visit physiotherapy practice. Built end to end, from the data model to the store build.

    FlutterDart
  • Personal

    Agent workflow system

    A two-layer system for working with AI coding agents that doesn't lock me into one tool. Runbooks for machines, essentially.

    PythonLLM tooling

Stack

cloud

AWSEC2S3VPCIAMCloudWatchRoute 53Cloudflare

infrastructure

TerraformDockerKubernetesLinuxNginx

code

PythonBashGitCI/CD

practice

MonitoringAlertingIncident responseOn-callCost optimisation

Now

AI engineering, pointed at operations.

Retrieval over runbooks, agents that can execute them, and the reliability practice that has to sit around a model once it's in production. It maps cleanly onto work I already do — a model endpoint is an API with worse failure modes, and LLMOps is mostly SRE wearing a new hat.

Alongside it I'm working toward the AWS DevOps Engineer Professional, CKA, and Terraform Associate.

Contact

Infrastructure that needs building, watching, or rescuing?

Write to me at [email protected], or find me on LinkedIn.

Start a conversation