I’m Product Lead, AI at Poe / Quora. Here’s one way I work through a difficult idea:
start with a road trip.
San FranciscoA starting state.
CarmelA decision along the way.
Los AngelesThe longer-term destination.
A good stop. A better journey?
In my reinforcement-learning session, a toy road trip separates immediate satisfaction
from expected long-term reward. The useful question isn’t only “what pays off now?”
It’s what that choice makes possible later.
The next experiment combines helpfulness, safety, and truthfulness signals into a
reward function.
HelpfulnessSafetyTruthfulness
→
Normalize. Weight. Combine.
A conceptual sketch of the livestream. The demo’s classifiers are imperfect proxies, not
validated measures of these qualities.
The caveat in context ↗