The next reward
isn’t always
the best one.

I’m Product Lead, AI at Poe / Quora. Here’s one way I work through a difficult idea: start with a road trip.

Illustration of a coastal journey from San Francisco through Carmel to Los Angeles
San FranciscoA starting state.
CarmelA decision along the way.
Los AngelesThe longer-term destination.

A good stop. A better journey?

In my reinforcement-learning session, a toy road trip separates immediate satisfaction from expected long-term reward. The useful question isn’t only “what pays off now?” It’s what that choice makes possible later.

Watch the road-trip example, from 16:21 ↗

From road trips
to model behavior.

The next experiment combines helpfulness, safety, and truthfulness signals into a reward function.

HelpfulnessSafetyTruthfulness
Normalize.
Weight.
Combine.

A conceptual sketch of the livestream. The demo’s classifiers are imperfect proxies, not validated measures of these qualities. The caveat in context ↗