I build reliable AI systems with Python, LangGraph, RAG, tool-using agents, evaluation pipelines, and FastAPI. My work focuses on making LLM applications observable, testable, secure, and useful in production-like environments.
I am a 2026 B.Tech graduate and AI Engineering Fellow at Maven (AI Makerspace). I build applied AI systems across agent workflows, retrieval-augmented generation, memory, evaluation, observability, privacy boundaries, tool safety, and cost-aware execution.
I am open to Applied AI Engineer, LLM Engineer, AI Engineer, and AI Evaluation roles in India, with preference for Bengaluru, Hyderabad, Mumbai, Delhi-NCR, Pune, and Noida.
- Agent reliability: Built Agent Reliability OS with trace collection, runtime tool-risk policies, secret redaction, baseline-vs-protected evaluations, CI, and a live demo.
- Agent security: Built MCP Sentinel Lab with policy evaluation, tool-risk scoring, redaction, CI, and a live MCP security demo.
- AI coding agents: Built DevMind with six security-aware tools, persistent sessions, runtime metrics, plugin support, CI, and 156 offline tests.
- Context engineering: Built ContextOps Agent with typed memory, plan persistence, context compression, privacy review, measurable evaluations, and CI.
- Evaluation and data generation: Built OpenAI AutoData with challenger, solver, and judge agents, budget controls, fail-closed validation, auditable outputs, 13 regression tests, and CI.
- Agent Reliability OS live demo · Repository · CI run
- MCP Sentinel Lab live demo · Repository · CI run
Production-style reliability layer for tool-using LLM agents.
- Built with: Python, FastAPI, SQLite, Streamlit, GitHub Actions
- Demonstrates: tracing, runtime policy enforcement, secret redaction, baseline-vs-protected evaluation, API, dashboard, and CI
- Proof: Live demo · Repository
Runtime security gateway and evaluation bench for MCP and tool-using AI agents.
- Built with: Python, policy evaluation, risk scoring, redaction, OpenRouter, GitHub Actions
- Demonstrates: tool-risk analysis, security controls, explainable policy decisions, CI, and a browser-based demo
- Proof: Live demo · Repository
Terminal-native AI coding agent built with Python, LangGraph, and Claude.
- Demonstrates: six built-in tools, persistent sessions, runtime metrics, plugins, cross-platform support, and 156 offline tests
- Proof: Repository · CI
Context-engineering layer for long-horizon agents with typed memory, compression, and privacy review.
- Demonstrates: plan persistence, memory reconstruction, privacy boundaries, token-savings metrics, API, dashboard, and CI
- Proof: Repository · CI
Agentic RAG system for trustworthy enterprise data integration and schema matching.
- Demonstrates: adaptive routing, evidence-backed decisions, OpenAI explanations, precision/recall/F1 evaluation, API, dashboard, and CI
- Proof: Repository · CI
- Secure RepoPilot: Issue-to-PR coding agent with baseline verification, command guardrails, privacy auditing, API, dashboard, and CI.
- OpenAI AutoData: Budget-aware multi-agent pipeline for generating difficult research QA data with validation and regression tests.
- Corrective Agentic RAG Assistant: Adaptive CRAG assistant with query routing, corrective retrieval, hierarchical context, web fallback, and RAG metrics.
- MemoryOS Agent: Long-term memory agent with OpenAI API support, SQLite memory, lifecycle controls, and a Streamlit dashboard.
- arXiv Digest Agent: Stateful research-paper digest and grounded QA workflow using retrieval, ranking, PDF parsing, embeddings, and FAISS.
- SKXYWTF Observability Platform: AI tracing, evaluation, cost/latency tracking, regression alerts, FastAPI, Supabase, and Streamlit dashboard.
Languages: Python, SQL
LLM and Agent Systems: OpenAI API, Anthropic API, LangGraph, LangChain, tool calling, agent memory, prompt engineering
RAG and Evaluation: Vector search, embeddings, corrective RAG, adaptive routing, citation grounding, precision/recall/F1, regression testing
Backend and Applications: FastAPI, Streamlit, SQLite, REST APIs
Reliability and Security: Observability, tracing, policy enforcement, secret redaction, privacy boundaries, cost controls
Engineering: Git, GitHub Actions, Docker, pytest, CI/CD
- LangChain Learning Lab: Prompts, LCEL, structured outputs, RAG, embeddings, and model integrations.
- Agno Basics: Tools, memory, RAG, multi-agent teams, workflows, AgentOS, and FastAPI.
- Adaptive RAG CAG Project: Adaptive RAG and cache-augmented generation demo with tests.
- Local RAG with Ollama and ChromaDB: Privacy-focused local PDF question answering.
- LangGraph Ollama Chatbot: Stateful local chatbot with SQLite checkpoints and token streaming.
- AI Reddit Brand Monitor: Local sentiment, topic, urgency, and feedback analysis for Reddit mentions.
- Build the smallest reliable system that proves the idea.
- Test failure paths, not only happy paths.
- Make cost, state, and model behavior visible.
- Keep claims aligned with reproducible code and results.
- Document limitations clearly instead of overstating benchmark performance.
For Applied AI roles, technical collaboration, or feedback on agent reliability and evaluation, feel free to reach out.
Forked repositories and profile configuration are kept separate from the flagship project portfolio.
