A road trip. A reward function.
A live reinforcement-learning experiment: combine helpfulness, safety, and truthfulness into one reward function.
Watch on YouTube
1:18:12
Recordings & writing
A live reinforcement-learning experiment: combine helpfulness, safety, and truthfulness into one reward function.
Watch on YouTube
1:18:12
A face-filter app, some coffees with strangers, and the feedback I didn’t act on. Notes from a project that failed.
Read the original retrospectiveI took paper mockups to a local coffee shop and offered to buy people a coffee for a few minutes of feedback. Those conversations raised problems with the idea. I kept building.
Later, time went into login, password reset, and features around the app. The quality and speed of the filters needed that attention instead. They were the core product—and a recurring point of feedback.
The 2016 retrospective separates the failure into design, coding, and distribution. The Maskito source code was opened up alongside it.