Skip to content
Suyash Bhavsar — AI Engineer, PuneOpen to new roles

I build voice & LLMsystems that scale.

Production voice agents, LLM/RAG pipelines and the low-latency, event-driven backends behind them — FastAPI, Redis and async processing.

1000+
Daily user queries
5,000+
Customers reached
500+
Concurrent sessions
12
Tenant orgs on the platform
Profile

I'm an AI Engineer at Graas in Pune, building real-time voice AI and LLM/RAG systems.

I design full-duplex, low-latency voice agents on OpenAI Realtime and Gemini Live, and production LLM pipelines — hybrid retrieval, prompt caching and evaluation — behind a conversational AI system handling 1000+ daily queries.

My backend work is Python (FastAPI) and PHP (Laravel) with Redis, Celery and OpenSearch: async, event-driven services, multi-tenant access control and observability for live voice and chat traffic.

Location
Pune, India
Education
M.Sc. Computer Science — MIT World Peace University (2022 – 2024)
Credentials
UGC-NET (June 2024) and Maharashtra SET — Assistant Professor eligibility in Computer Science
Focus
Real-time voice AI, LLM/RAG, AI agents & distributed backend systems
Availability
Open to AI engineering roles — Pune or remote (India).
Suyash Bhavsar — portrait
Suyash BhavsarPune · 18.52°N 73.86°E

What I do

Real-time voice AI

Full-duplex voice agents over WebRTC and WebSocket with a pluggable provider layer (OpenAI Realtime, Gemini Live) and live-agent hand-off.

LLM & RAG systems

Hybrid retrieval — structured filtering, KNN semantic search and keyword fallback — with prompt caching and manual-plus-A/B evaluation.

Async backend engineering

FastAPI and Laravel services on Redis-backed queues and gevent-pooled Celery workers, built for 500+ concurrent sessions.

Multi-tenant platforms

OTP auth, org-scoped data isolation and a hook-based customization framework used by 10+ client organizations.

Stack

Tools I ship with

Voice AI & LLM

OpenAI Realtime APIGemini Live APILLMsRAGHybrid SearchPrompt EngineeringLangChainMCPPineconeLLM EvaluationTensorFlowKeras

Backend & Real-Time

PythonFastAPIPHP (Laravel)asyncioFlaskSpring BootREST APIsMicroservicesMulti-Tenant ArchitectureWebRTCWebSocketsRedisCelery

Cloud, Data & Observability

AWS (CodeDeploy)DockerPostgreSQLOpenSearchMySQLMongoDBSQLGoogle Chat Alerting

Tools, Testing & Frontend

GitLinuxPytestPestGitHubPostmanJiraClaude CodeReactTailwind CSSViteAlpine.jsJavaScript

Experience

Where I've worked

  1. Jul 2024 — Present

    Pune, India

    AI Engineer — Graas

    • Designed, built and deployed Graas's real-time voice AI system end to end for WhatsApp voice calls and web — a full-duplex WebRTC/WebSocket pipeline with a pluggable provider layer (OpenAI Realtime, Gemini Live) and live-agent hand-off; scaled to outbound campaigns reaching 5,000+ customers.
    • Built live call-quality and latency monitoring that brings root-cause debugging of production voice issues to under 14 minutes, plus an evaluation system combining manual review and A/B testing.
    • Developed an LLM-powered conversational AI system handling 1000+ daily user queries, with hybrid product retrieval (DSL-based filtering, KNN semantic search, structured → vector → keyword fallback).
    • Cut response latency by ~40% with Redis-based asynchronous processing — queue-based webhook ingestion decoupled from LLM generation, per-organization prompt caching and short-TTL product caching — on FastAPI and Laravel services supporting 500+ concurrent sessions.
    • Engineered a versioned hook-based customization framework with a UI function runner, used by 10+ client organizations to add custom tools and scheduled jobs without code changes or redeploys.
    • Integrated WhatsApp and Facebook APIs (70% of interactions via messaging channels) and designed multi-tenant access control — OTP auth and org-scoped data isolation across 12 tenant organizations.
  2. Jan 2024 — Jun 2024

    Pune, India

    Software Engineer Intern — Graas

    • Built a Spring Boot data-operations automation framework for metrics query workflows with approval tracking, reducing manual verification steps across 5 teams.
    • Developed a static impact-analysis tool for Magento using PHPStan, PHPParser and AST parsing; it cut manual regression-testing time by about 30% across 50+ modules.
Selected work

Built, tested, shipped

Voice Calling Agent preview

Voice AI · Proof of concept

Voice Calling Agent

Real-time streaming AI voice system

A proof-of-concept voice agent: the browser captures microphone audio and opens a WebRTC session with a FastAPI backend, which streams it to an AI provider (OpenAI or Gemini) and plays the voice reply back with live transcripts over WebSocket. A pluggable provider layer sits behind a unified interface, with mid-conversation tool calling (mocked in the repo) and handling of partial transcripts and interruptions. Per-call latency instrumentation logs each stage from WebRTC to AI response to tool call; the repo's tool calls run against a SQLite product catalogue.

Full-duplex · per-stage latency logging

Python / FastAPI / WebRTC / WebSockets / OpenAI / Gemini

AICompete preview

RAG · Backend

AICompete

AI competitive intelligence platform

An end-to-end competitive intelligence platform built on FastAPI, Celery and scraping pipelines, with LLM-powered summarization and a RAG pipeline over Pinecone embeddings. Background analytics and alerting workflows on Scrapy, Redis and PostgreSQL drive sentiment tracking, keyword detection and automated competitor monitoring.

FastAPI / Celery / Scrapy / Redis / PostgreSQL / Pinecone / LangChain

Tailwind Magic preview

Developer Tooling

Tailwind Magic

CSS-to-Tailwind converter and refactoring utility

A command-line utility that converts legacy CSS into Tailwind utility classes across HTML, React, Vue and Angular projects. It uses AST-based parsing and a reverse lookup of Tailwind styles, maps media queries to responsive prefixes, and offers automatic backups, a dry-run mode and detailed conversion reports.

Node.js / JavaScript / AST parsing / Tailwind CSS

GPAVBHAG preview

Web App

GPAVBHAG

Restaurant website

A responsive restaurant web application with menu browsing and item pages, a validated contact form, and an interactive location map. Built with React Router and lazy-loaded pages, Context API state, custom hooks for forms, localStorage and weather data, and a mobile-first Tailwind CSS layout.

React 19 / Vite / Tailwind CSS / React Router

Handwriting Classifier preview

Machine Learning

Handwriting Classifier

Human vs. computer-generated text

A web app that classifies an image of a word as human-written or computer-generated using a CNN. It trains on handwritten words plus a generated dataset of words rendered in many fonts, with mixed-precision training, logging, and a Flask interface for image upload and prediction.

Python / TensorFlow / Keras / Flask

Spell Checker preview

Web App

Spell Checker

Real-time spelling suggestions

A spell-checking web app that flags misspellings as you type, offers live suggestions, and lets you add words to a custom dictionary or clear it. Built with Laravel, Alpine.js and Tailwind CSS.

Laravel / Alpine.js / Tailwind CSS

Contact

Have something worth shipping?

Tell me about the product, the problem, or the role — email works best.

bhavsarsuyash30@gmail.com

Open to AI engineering roles — Pune or remote (India).