Production AI
Generative AI, RAG, agents, NLP-to-SQL, fine-tuning, evaluation, and real-time inference.
I build production AI systems using agents, LLMs, NLP, computer vision, and deep learning. My focus is simple: good accuracy, low latency, and reliable performance.
I connect applied AI research with disciplined software engineering to deliver systems that are accurate, observable, scalable, and useful in production.
Generative AI, RAG, agents, NLP-to-SQL, fine-tuning, evaluation, and real-time inference.
From training pipelines and APIs to event-driven AWS architecture, Kubernetes, and observability.
Accuracy, p90 latency, concurrency, reliability, and cost are treated as core product requirements.
Led teams of 10+ members, completed 40+ code reviews, and contributed across design, delivery, and deployment.
Designed production RAG systems, intelligent tool-routing agents, memory-enabled workflows, and real-time inference services.
Built 100+ production APIs following software-engineering standards and integrated AI capabilities into new and existing applications.
Optimized orchestration accuracy, NLP-to-SQL, retrieval quality, model inference, concurrency, and end-to-end latency.
Migrated deep-learning training and inference workloads from NVIDIA GPU infrastructure to AWS Trainium and Inferentia.
Worked across collaboration, audio generation, solar compliance, multilingual translation, and manufacturing quality domains.
Led 10+ team members, completed 40+ code reviews, and took ownership across architecture, implementation, optimization, and deployment.
Five production-focused engagements spanning agents, deep learning infrastructure, computer vision, multilingual inference, and industrial automation.
Memory-enabled enterprise messaging intelligence
An agentic assistant for field workers that retrieves project, people, meeting, outing, and date-based information from PostgreSQL messaging data with low latency.
Audio generation on AWS Trainium & Inferentia
Migration of a multi-model audio-generation training and inference stack from CUDA/B200 GPUs to AWS purpose-built AI accelerators.
Automated solar-installation compliance
An AI quality-control platform that evaluates installation photos, identifies defects, references compliance knowledge, and returns image-level pass/fail guidance.
Enterprise real-time translation across 76 languages
A multilingual inference system designed around model-quality routing, large-scale caching, and a p90 real-time latency target.
Real-time manufacturing quality intelligence
Automated visual inspection for a factory producing 20,000 PCB boards each day across ten production lines.
Saveetha Engineering College
Professional-level machine learning engineering credential.
Reach out for enterprise AI engagements, consulting, or collaboration.