Date listed
5 days agoEmployment Type
Full timeFound on:
We work with unusual intensity. In-person in San Francisco, six days a week, long days, most weekends. This is not a phase we'll grow out of. It's how we've chosen to build, because we're in a market where speed decides who wins.
We're telling you this in the first paragraph, not the last, because we only want people who read that and feel pulled in, not talked into it. If you want a 9-to-5 (genuinely, no judgment), this isn't your role, and we'd rather you know now.
Here's what you get in exchange:
Cekura (YC F24) is building the voice AI engineer. Teams use Cekura to test agents before going live, monitor real production calls, and self-improve continuously. Cekura doesn't just flag issues and suggest fixes: it reproduces failures in simulation, fixes them, tests the fix thoroughly, and raises PRs. The platform spans pre-production simulation, LLM-powered evaluation, adversarial red-teaming, production monitoring with live drift detection, and cross-provider benchmarking (Vapi, Retell, Pipecat, LiveKit, ElevenLabs, and more).
We're trusted where reliability is non-negotiable: customers include Five9, HighLevel, Twin Health, PwC, Deloitte, Jobber, and Jotform, with HIPAA, SOC 2, and GDPR as defaults, not checkboxes. We're growing fast and backed by top investors.
You'll build the core of Cekura: the simulation engines, evaluation systems, self-improvement loops, and observability pipelines our customers rely on to ship voice agents with confidence.
We deliberately don't split this into "software engineer" vs. "AI engineer." The interesting problems live at the boundary: real-time voice infrastructure meets LLM-as-judge evaluation, distributed systems meet RL-style self-improvement loops, telephony meets audio and speech analysis (ASR quality, barge-in, latency, prosody), and classic NLP meets frontier agentic behavior. You'll work across that whole surface.
Build the testing and simulation engine. Design and ship systems that simulate thousands of realistic conversations against customer agents across voice, chat, and phone, with control over personas, interruptions, background noise, and edge cases.
Push the frontier of agent evaluation. Build LLM-powered evaluators, metrics, and the closed self-improvement loop at the heart of the product: detect failure, reproduce in simulation, generate the fix, test it thoroughly, and raise the PR, all autonomously. This includes adversarial red-teaming for jailbreaks, PII leaks, and off-script behavior, plus production monitoring with live drift detection.
Do applied audio and speech research. Go beyond the transcript. Separate background noise from actual speech, and map the paralinguistic layer of conversation (emotion, tone, silences, hesitations, speaking rate, overlaps) into structured signals, so agents are evaluated on how something was said, not just what.
Own real-time voice infrastructure. SIP, WebRTC, WebSockets, STT/TTS pipelines, and providers like Twilio, Vapi, Retell, LiveKit, and Pipecat. Latency, barge-in, and audio quality are first-class problems here.
Ship end-to-end. Take features from design to production. You own the full stack of what you build: backend, infra, evals, and the product surface customers touch.
Shape how we build. We're a small, senior team. Your architectural decisions, code standards, and technical taste will compound as the team grows.
Newsletter
Let's simplify your job search. Receive your tailored set of opportunities today.
Subscribe to our Jobs