All writing

Gate B4Writing

RICE breaks a little when the feature is an LLM

Reach and confidence stop meaning what they used to. Here is the version of RICE I actually use for AI features.

RICE was written for a world where features are deterministic. You ship the button, users click it or they don't, and reach is measurable within a week.

AI features don't behave that way. The feature works, then it works differently, then a model update changes what "works" means. I've been adjusting each of the four levers.

Reach

Standard reach — "users per quarter" — is still fine as a headline. But for AI features I add a second number: eligible reach. How many of those users will actually hit the feature in a way where the model responds well?

Impact

The 3 / 2 / 1 / 0.5 / 0.25 scale is fine, but I write down the expected impact and the 90th percentile impact separately. AI features have a heavier tail — a chat feature might land at "moderate" on average but "huge" for the users whose task the model handles well. Pricing both matters.

Confidence

This is where RICE gets hardest for AI. Confidence is normally 100/80/50%. For AI features I break it into three parts:

  • Product confidence — will users want this?
  • Eval confidence — does the model do this well enough today?
  • Ops confidence — can we run this in production without breaking cost or latency SLAs?

I multiply them. A feature where I'm 80% sure users want it, 60% sure the model can do it today, and 90% sure we can run it, lands at 43% — not 80%.

Effort

Effort for AI features is bimodal. Either it's "prompt engineering plus a bit of UI" (small), or it's "we need a real eval harness, monitoring, and a rollback plan" (large). I've stopped estimating in the middle. If the feature needs a real eval loop, I write it down at 3x what I'd normally estimate.

What I don't do

  • I don't include the model itself in effort. That belongs on the ML roadmap, not the product roadmap.
  • I don't compare AI and non-AI features on the same sheet without separating them by their confidence formulas. Otherwise the AI features always look worse and get bumped.