Applied AI Engineer San Francisco, CA

Chandhan Saai Katuri

Production LLM systems Open to new roles

I'm a software engineer working on applied AI. I build the production systems that sit around machine learning models: the pipelines that feed them, the evaluation that checks their output, and the infrastructure that keeps them running.

About

My background is in robotics. I spent my master's at the University of Maryland working on perception and control, where a system that is confidently wrong does not fail quietly. That experience shaped how I approach machine learning: I am less interested in whether a system produces an answer than in whether we can tell when the answer is wrong.

I work across the full stack, largely because the problems worth solving rarely stay in one layer. A slow page turns out to be a missing database index; a model that appears accurate turns out to be a pipeline that stopped running. Diagnosis is the part of the work I find most rewarding.

I place particular value on the engineering that is easy to skip: tests that genuinely fail when the code breaks, alerting that fires when it should, and migrations that can be rolled back safely.

Experience

Current Apr 2026 — Full breakdown →

Verita AI Software Engineer

Designed and shipped a vision-LLM audit pipeline that reads contractor screen recordings against per-project task policies and seals priced, recomputable audit ledgers.

Holds up 80% precision and recall on the golden evaluation set; 264 sessions sealed in production. Read models benchmarked across providers before one was chosen.

Architected the core of an LLM data-annotation platform: five-role access control, an audited workflow state machine, an append-only audit log.

Holds up Audit coverage proven across all 54 endpoints; test suite grew 210 → 1,239. Django REST Framework and React.

Delivered a finance-interview rubric verifier in six days: LLM-judge claim validation, shadow-verifier gating, hard execution limits.

Holds up Set the team standard that a test only counts once a deliberate source defect proves it can fail, and the defect itself is verified as landed.

Led SOC 2 Type II readiness during an active audit window, including a pgAudit and storage-encryption migration across 11 production Postgres databases.

Holds up Control coverage moved 154 → 165 of 173 (95%), with 40+ evidence artifacts and penetration-test findings remediated.

Built a multi-scanner security pipeline with LLM-assisted triage across five static-analysis engines.

Holds up 345 raw findings reduced to 33 verified, nine of them critical: remote code execution, SSRF, IDOR, exposed credentials. Pipeline is public on GitHub.

Closed Sep 2025 – May 2026 Full breakdown →

Handshake AI AI Data & Evaluation

Evaluated frontier image-generation models against calibrated quality rubrics to produce preference data for RLHF pipelines.

Holds up Rubrics covered rendering artifacts and anatomical consistency; also reviewed ML codebases for algorithm selection and implementation quality.

Selected as a Star Fellow on a program probing model reasoning boundaries.

Holds up The work ran against frontier research publications, testing where model reasoning breaks down.

Open source

Stack

LLM systems

  • Vision-model pipelines
  • RAG
  • FAISS
  • Pinecone
  • LangChain
  • LangGraph
  • LLM-as-judge
  • Golden datasets
  • Mutation testing
  • OpenAI
  • Anthropic
  • OpenRouter

Backend

  • Python
  • Django
  • DRF
  • FastAPI
  • PostgreSQL
  • Celery
  • SQS
  • REST APIs

Frontend

  • TypeScript
  • React
  • Vite
  • Vitest

Infrastructure

  • AWS
  • S3
  • RDS
  • Lambda
  • CloudFront
  • CloudWatch
  • GuardDuty
  • Terraform
  • Docker
  • GitHub Actions
  • Sentry

Machine learning

  • PyTorch
  • Transformers
  • YOLO
  • ViT
  • ONNX
  • TensorRT

Education

  • M.Eng Robotics, University of Maryland
  • B.Tech Mechatronics, SRM Institute of Science and Technology

Contact me

Email me GitHub LinkedIn