Senior Voice AI Engineer at Birdeye in Gurugram, India - Apply now!
R&DFull-timedescription
About the Role
We are looking for a Senior Voice AI Engineer to build and scale next-generation conversational Voice AI systems. You will work on real-time, low-latency voice pipelines involving Speech-to-Text (STT), Large Language Models (LLMs), Text-to-Speech (TTS), WebRTC, telephony, and streaming infrastructure.
You should have hands-on experience building production-grade Voice AI agents using frameworks like Pipecat, LiveKit, or similar real-time communication platforms. This role requires deep understanding of streaming architectures, WebSockets, voice quality optimization, interruption handling, and scalable backend systems.
About Birdeye
Birdeye is the leading agentic marketing platform for multi-location brands.
Companies like H&R Block, Aspen Dental, and Caesars Entertainment use Birdeye to manage marketing across thousands of locations — from how they get found, to how they convert, to how they retain customers. Our platform replaces disconnected point tools with AI agents that execute work at the location level — responding to reviews, updating listings, publishing content, and driving conversions.
Backed by Marc Benioff, Jerry Yang, and Accel-KKR, Birdeye was named to G2’s 2026 Best Agentic AI Products list — appearing alongside the world’s leading AI companies. We’re expanding rapidly into enterprise, with growing adoption across large, multi-location brands.
Responsibilities:
- Design and build production-grade Voice AI applications.
- Develop real-time streaming pipelines for voice conversations.
- Integrate STT, LLM, and TTS providers into low-latency conversational systems.
- Build robust WebSocket/WebRTC infrastructure for bi-directional audio streaming.
- Implement interruption handling (barge-in), turn detection, and Voice Activity Detection (VAD).
- Optimize latency across the complete voice pipeline.
- Integrate telephony providers such as Twilio, SIP, or LiveKit Telephony.
- Develop scalable backend services capable of supporting thousands of concurrent voice sessions.
- Design session management, state management, and conversation memory.
- Build monitoring, observability, and analytics for voice conversations.
- Deploy and operate Voice AI infrastructure on Kubernetes and cloud platforms.
Required Skills:
Voice AI
- Strong understanding of conversational Voice AI systems
- Experience building real-time AI voice assistants
- Knowledge of latency optimization techniques
- Understanding of conversational memory and dialogue management
Speech Technologies
- Experience with OpenAI Realtime API, ElevenLabs, Deepgram, AssemblyAI, Google Speech, Azure Speech, Amazon Transcribe, Cartesia, or PlayHT
- Streaming STT and TTS
- Voice cloning and adaptive speech
- Speaker diarization
- Custom vocabulary and pronunciation dictionaries
Real-Time Communication
- Pipecat
- LiveKit
- WebRTC
- RTP
- SIP
- WebSockets
- Server-Sent Events (SSE)
Backend Development
- Python (preferred), FastAPI, AsyncIO
- WebSocket servers, gRPC, REST APIs
- Event-driven architecture
- Redis, Kafka, RabbitMQ
- PostgreSQL, MongoDB, ClickHouse (good to have)
AI & LLM
- Experience with OpenAI, Anthropic, Google Gemini, or open-source LLMs
- Prompt engineering, tool calling, function calling
- RAG, agentic workflows, multi-agent orchestration, conversation memory
Cloud & Infrastructure
- AWS, GCP, or Azure
- Docker
- Kubernetes
- NGINX
- Load Balancers
- Autoscaling
- CI/CD
requirements
Preferred Experience:
- Built production Voice AI products
- Experience with Pipecat in production
- Experience with LiveKit Cloud or self-hosted LiveKit
- Experience with Twilio Voice or SIP infrastructure
- AI phone agents and call center automation
- Noise suppression (Krisp, DeepFilterNet, RNNoise)
- Experience supporting 1,000+ concurrent voice sessions
- Multilingual Voice AI
Qualifications:
- Bachelor's or Master's degree in Computer Science or related field.
- 6-8 years of software engineering experience.
- 3+ years building real-time streaming systems.
- 2+ years working on Voice AI or conversational AI products.
- Strong system design and distributed systems experience.
Why You'll Join Us:
At Birdeye, we are relentless innovators driven by a singular goal: to lead our category with unparalleled excellence. We don't just set goals – we surpass them. We're a team of doers who roll up our sleeves and get the job done, delivering on our promises with unwavering dedication.
Working here means embracing a culture of action and accountability, where every person is empowered to make an impact. We don't just talk about making a difference – we make it happen.