Cyanotype portrait of Gareth explaining an idea with his hands

What counts
as helpful?

Product Lead, AI · Poe / Quora

A model can sound helpful without being right. I’m interested in how we turn that uncomfortable gap into something we can test.

Three signals.
No perfect judge.

A sketch of the reward-function experiment—not a scorecard or a validated measure of safety.

Helpfulness

The demo uses sentiment as a proxy.

Safety

A toxicity classifier supplies another signal.

Truthfulness

A fake-news classifier adds an imperfect third view.

Normalize + weight the signals→ Combined reward

Adapted from the livestream. These proxy choices are part of the experiment, not proof of the qualities they stand for.

First, a road trip.

The session starts smaller: a journey from San Francisco through Carmel to LA, and the difference between the next reward and the long-term one.

Start with the teaching example ↗

Then, real tasks.

My agent-evaluation session moves the question into live environments: what can an agent actually get done?

Evals for AI Agents ↗