Evals for AI Agents
What happens when agents meet real tasks in live environments?
Watch the evaluation workshop ↗A reward function has to hold several ideas at once. In this experiment: helpfulness, safety, and truthfulness.
The demo uses sentiment analysis as a proxy. Positive sentiment is not the same as a helpful answer—a useful limitation to keep visible.
Adapted from the Multi-Signal Reward Function workshop. These are experimental proxies, not validated measures or production safety guarantees.
Watch the full experiment · 1:18:12 ↗What happens when agents meet real tasks in live environments?
Watch the evaluation workshop ↗A simpler route from a spoken idea to text at your cursor.
Explore the project ↗