AI Engineer | Chennai, India

I build production AI systems using agents, LLMs, NLP, computer vision, and deep learning. My focus is simple: good accuracy, low latency, and reliable performance.

0%+
Agent orchestration
0%
NLP-to-SQL accuracy
<0.0s
Real-time AI target
0+
Production APIs
Jagan Sivakumaran, AI Engineer
AI Engineer
Chennai | Remote friendly
AGENTIC AILLM SYSTEMSCOMPUTER VISIONREAL-TIME INFERENCEAWS AI/MLDEEP LEARNINGAGENTIC AILLM SYSTEMSCOMPUTER VISIONREAL-TIME INFERENCEAWS AI/MLDEEP LEARNING
Profile

Engineering AI that performs outside the demo.

I connect applied AI research with disciplined software engineering to deliver systems that are accurate, observable, scalable, and useful in production.

Production AI

Generative AI, RAG, agents, NLP-to-SQL, fine-tuning, evaluation, and real-time inference.

End-to-end engineering

From training pipelines and APIs to event-driven AWS architecture, Kubernetes, and observability.

Measured outcomes

Accuracy, p90 latency, concurrency, reliability, and cost are treated as core product requirements.

Technical ownership

Led teams of 10+ members, completed 40+ code reviews, and contributed across design, delivery, and deployment.

Professional experience

AI Engineer | 1.5+ years

Designed production RAG systems, intelligent tool-routing agents, memory-enabled workflows, and real-time inference services.

Built 100+ production APIs following software-engineering standards and integrated AI capabilities into new and existing applications.

Optimized orchestration accuracy, NLP-to-SQL, retrieval quality, model inference, concurrency, and end-to-end latency.

Migrated deep-learning training and inference workloads from NVIDIA GPU infrastructure to AWS Trainium and Inferentia.

Worked across collaboration, audio generation, solar compliance, multilingual translation, and manufacturing quality domains.

Led 10+ team members, completed 40+ code reviews, and took ownership across architecture, implementation, optimization, and deployment.

Selected work

Complex AI systems, made operational.

Five production-focused engagements spanning agents, deep learning infrastructure, computer vision, multilingual inference, and industrial automation.

PROJECT 01

RedeAPP

Memory-enabled enterprise messaging intelligence

An agentic assistant for field workers that retrieves project, people, meeting, outing, and date-based information from PostgreSQL messaging data with low latency.

StrandsClaude Sonnet 4.5BedrockAgentCorePostgreSQLAWS Knowledge Bases
01Strands orchestration routes semantic questions to AWS Knowledge Bases and date/data questions to an NLP-to-SQL tool.
02A scheduled two-day preprocessing pipeline embeds recent data with Amazon Titan Embeddings V1 and stores vectors through AWS Knowledge Bases.
03AgentCore Memory provides long- and short-term conversation memory; AgentCore Runtime hosts the production agent.
95%+ orchestration accuracy
98% NLP-to-SQL accuracy
100% RAG accuracy
<1,200 ms at 50 concurrent users
PROJECT 02

SoundryAI

Audio generation on AWS Trainium & Inferentia

Migration of a multi-model audio-generation training and inference stack from CUDA/B200 GPUs to AWS purpose-built AI accelerators.

AWS NeuronTrainiumInferentiaEKSPyTorchHTDemucsFlowVAERVQ
01Migrated training to trn2.48xlarge and inference to inf2.24xlarge, using tensor parallelism across NeuronCores to resolve memory constraints.
02Stabilized training with memory optimization and gradient clipping, then built a UI to configure and launch training jobs.
03Deployed inference on EKS with HTDemucs, an autoregressive generator, RVQ tokenization, and FlowVAE decoding; PostgreSQL notifications queue requests for a five-user workload.
~2 min for a 3 min song
TP across 8 NeuronCores
Trainium training
Inferentia inference
PROJECT 03

Eequals Solar QC

Automated solar-installation compliance

An AI quality-control platform that evaluates installation photos, identifies defects, references compliance knowledge, and returns image-level pass/fail guidance.

RF-DETROCRLangChainLangGraphRAGAWS
01Fine-tuned RF-DETR to detect installation defects and combined it with OCR for labels and component evidence.
02Used LangChain for agent tooling and LangGraph for routing across visual analysis and a modifiable compliance RAG knowledge base.
03Designed low-confidence escalation, partial re-evaluation, admin-controlled knowledge updates, and AWS environments for Dev, UAT, and Production.
Seconds instead of hours
Per-image verdicts
Human escalation path
Designed for 15k images/month
PROJECT 04

Airbnb Translation

Enterprise real-time translation across 76 languages

A multilingual inference system designed around model-quality routing, large-scale caching, and a p90 real-time latency target.

TranslateGemma 27BRLHFvLLMB200RedisBedrockEKS
01Benchmarked models across 76 x 76 language pairs on eight B200 GPUs and selected TranslateGemma 27B with Claude Sonnet 4.5 fallback.
02Built a request path from a 1.5 TB Redis cache to accuracy-metric JSON routing, then to the best available translation model.
03Prepared RLHF data with Claude Opus 4.5, implemented reward-based training with model parallelism and vLLM forward passes, and deployed inference on EKS.
95% model accuracy
76 languages
1.5 TB Redis cache
1,500 ms p90 delivered
PROJECT 05

AI-Powered PCB Defect Inspection

Real-time manufacturing quality intelligence

Automated visual inspection for a factory producing 20,000 PCB boards each day across ten production lines.

Computer VisionSFTPydantic AIPostgreSQLVector SearchPython
01Built a ten-class defect detector using client and externally sourced data, expanded through data augmentation and supervised fine-tuning.
02Triggered a Pydantic AI agent after detection, passing the image and predicted class into a remediation workflow.
03Stored known-error knowledge in PostgreSQL with vector search so the agent can recommend line stops, machine routing, or corrective action.
96% defect accuracy
20k boards/day
10 production lines
10 defect classes
Technical toolkit

Breadth across the AI product stack.

AI & ML

Deep LearningMachine LearningNLPLLMsComputer VisionRAGSFTRLHFPrompt Engineering

Agentic AI

StrandsPydantic AILangChainLangGraphLangSmithAWS AgentCore

Backend & Data

PythonFastAPISQLNoSQLPostgreSQLMongoDBMySQLVector DatabasesWebSocket

Cloud & MLOps

AWS LambdaS3SQSECSEKSEC2BedrockSageMaker AIKubeflowDatabricksDockerKubernetesCI/CD

Frontend & Quality

ReactTypeScriptJavaScriptHTMLCSSPostmanThunder ClientTesting
Credentials

Education & certification.

2021 - 2025 | Chennai, India

B.E. Computer Science

Saveetha Engineering College

Professional certification

Databricks Certified Machine Learning Professional

Professional-level machine learning engineering credential.

Contact

Let's build something intelligent.

Reach out for enterprise AI engagements, consulting, or collaboration.

Direct
jagansivakumaran@gmail.com
Response Time< 4 hours
LocationChennai, India
SpecializationEnterprise AI systems