We companies build and launch public benchmarks that measure what AI can do in real-world domains. Our work includes ITSMBench, which evaluates whether AI models can autonomously resolve IT tickets. We’re looking for a intern to help build domain-specific benchmarks, develop tools that make benchmarking easier, and research ways to create evaluations that are fair, challenging, and useful.
What you’ll work on
What we’re looking for
Experience with evaluation systems or research projects is a plus, but isn’t required.
There is a potential opportunity to join full time after three months.
Senior Manager Software Engineering , Twilio
2 months ago
Tech leads (Platform and Software) and Platform engineers , EggAI
1 month ago
Founding Engineer , Dave Evans
| $170k - $210k1 month ago
Engineering Manager, Applied AI , JUPUS
4 days ago
Full Stack Engineer , Entangl
Full Time Employment2 weeks ago
Newsletter
Let's simplify your job search. Receive your tailored set of opportunities today.
Subscribe to our Jobs