About
I’m Murali, a senior engineer working across AI engineering and platform engineering — I build agentic AI systems and the production platforms that run them: reliable, observable, secure, and cost-aware.
My work sits at the intersection of AI engineering and platform engineering. I work across the full lifecycle of an agent system: how agents reason and collaborate, how they use tools and retrieve context, where they run, how their quality is measured, and how the platform is operated safely at scale.
What I’m working on
In my current role, I build AI platforms that combine large language models with mathematical optimization and operations research. The agent layer includes multi-agent orchestration with the Claude Agent SDK and LangGraph, specialized agents, memory and context management, retrieval-augmented generation (RAG), and tool integrations exposed through MCP servers built with FastMCP.
The runtime layer spans FastAPI services, AWS ECS and Fargate, GCP Cloud Run, Cloudflare, and isolated Modal sandboxes. I build provider-agnostic model routing across Amazon Bedrock and Google Vertex AI so workloads can use the right model and region while meeting EU data-residency and GDPR requirements.
Getting an agent to produce a good answer once is only the beginning. I build evaluation pipelines around golden datasets, synthetic test variants, quantitative thresholds, and LLM-as-judge scoring through Claude on Amazon Bedrock. These evaluations run as CI/CD quality gates in GitHub Actions, making regressions visible before they reach production. On the operational side, I work on agent-level observability, application monitoring, token and latency tracking, and per-tenant usage and cost attribution using services such as CloudWatch, DynamoDB, Athena, and QuickSight.
Underneath the agents, I design and automate the platform with Terraform, GitHub Actions, and environment-specific deployment pipelines. That includes secure runtimes, blue-green deployments, identity and tenant-scoped authorization with Amazon Cognito and JWT, secrets management, monitoring, and the AWS and GCP infrastructure needed to run the systems reliably.
The experience behind the platform
I bring more than a decade of experience across cloud infrastructure, security, SRE, and platform engineering. Before focusing on agentic AI, I built internal RAG and natural-language analytics applications with Amazon Bedrock, FastAPI services on AWS, reusable Terraform modules, data platforms with Amazon Redshift and DMS, and production observability with Datadog.
Earlier in my career, I worked across AWS, GCP, and Azure on Kubernetes and container platforms including EKS, GKE, and ECS; production secrets management with HashiCorp Vault; IAM and federated identity; DevSecOps automation; incident response; and cloud security. That background shapes how I approach AI systems today: the model matters, but so do the runtime, permissions, failure modes, delivery pipeline, telemetry, and cost model around it.
I’m continuing to deepen my expertise across AI engineering and AI platform engineering. I learn by building, and I write here about the systems I’m working on, the problems I run into, and what I learn while solving them.
Certifications
AWS Solutions Architect Professional · DevOps Engineer Professional · Security Specialty · ML Engineer Associate · Data Engineer Associate · SysOps Administrator · Developer Associate · Solutions Architect Associate · AI Practitioner · Cloud Practitioner
Kubernetes CKA — Certified Kubernetes Administrator · CKS — Certified Kubernetes Security Specialist
HashiCorp · GCP · Azure Vault Associate · GCP Associate Cloud Engineer · Azure Fundamentals
All badges verified on Credly →
LinkedIn · GitHub · X · muralidkt@gmail.com