Date listed
3 days agoEmployment Type
Full timeFound on:
Lemma is production monitoring for AI agents. We catch the silent failures your observability tools and evals miss (think bad tool calls, lost context, and infinite loops) before your users find them.
Agents break silently. They call the wrong tool, forget what the user said three turns ago, and loop until someone pulls the plug. Soon they'll be responsible for the majority of the world’s economic work, and most teams won't even know when they fail.
Making agents reliable is the problem Lemma exists to solve. That means catching the unknown unknowns, the failures nobody thought to write an eval for, and closing the loop end to end so they get fixed, not just flagged. It's the foundation for building agents people can actually trust and the future of self-improving systems.
The hardest part of our product is deciding what counts as a failure.
There is no ground truth here, and no benchmark to climb. Every customer's agent is different, what "wrong" means changes from one to the next, and we have to get it right across production without anyone telling us what to look for. Being confidently wrong often costs us more trust than being right fifty times earns.
This role owns the intelligence in the loop: what we flag, how sure we are, and whether the fix we propose actually fixes it.
You'll join a team of dropout founders and engineers from Amazon, Together, and Zoom. We've been founding operators at unicorns and at startups that went on to be acquired.
Onsite in San Francisco. We don’t sponsor visas.
Newsletter
Let's simplify your job search. Receive your tailored set of opportunities today.
Subscribe to our Jobs