back to writing

What makes an AI product stick

·A practical way to evaluate AI products beyond technical capability

I moved to San Francisco recently to build an AI startup. You can't go two minutes here without a mention of AI, but building one myself made me think harder about what actually sticks.

When a new AI product launches, the conversation usually starts with capability. What can it do? How good is the model? How many tasks can it handle?

That's not where I start.

Most AI tools today are functionally similar on paper. The differences that matter only appear when you use them for real work. Two products can share the same underlying model and feel completely different in practice.

What I actually look for is behaviour under live conditions: how a product handles ambiguity, failure, iteration, and the constraints of real workflows. A few signals consistently separate the thoughtfully built products from the ones that are merely impressive on first use.

1. How the product handles ambiguity

Most real tasks are underspecified. People arrive with half-formed goals, incomplete context, and constraints they haven’t fully worked out yet.

One of the first things I test is what happens when I'm vague. Does the product ask a clarifying question? Make a reasonable assumption and state it? Or does it confidently go in the wrong direction?

Products that take ambiguity seriously feel more like collaborators. The ones that ignore it create extra work, even when the initial output looks good.

2. Where friction is a design choice

Not all friction is bad. In AI systems, friction is often how risk gets managed.

I pay attention to where a product introduces pauses, confirmations, or review steps. These slow the user down just enough to prevent costly mistakes. Removing that friction can speed up low-stakes work. In higher-stakes contexts, the same design choice can quietly transfer risk to the user.

These decisions have nothing to do with model quality. They reflect product judgment in how the team thinks about trust, reversibility, and who takes responsibility when something goes wrong at scale.

3. Whether the product helps you recover

I rarely judge a product by its first output. What I care about is how it behaves when something is wrong, unclear, or incomplete.

How easy is it to recover? Can I course-correct without starting over? Does the product preserve useful context, or lose it? Does it acknowledge a mistake, or double down on one?

Resilient products make iteration easy. Brittle ones break after a single bad interaction. How a product recovers tells you just as much about how it performs.

4. How well the product fits existing workflows

I don't use AI tools in isolation. They live alongside documents, calendars, messaging apps, and task systems.

When I compare products, I look at whether a tool forces a new workflow or fits into an existing one, and how much context switching it demands. Tools that want to become the center of your workflow usually aren't worth the reorganization.

The best products feel boring in the right way. They surface at the moment of need, work with existing artifacts, and disappear when they're done.

Take Granola. It sits in the background during a meeting, captures what's said, and produces notes worth reading, all without changing how you run the meeting. You don't notice it while it's working. That's the point!

5. Whether the product earns daily use

After twenty or thirty minutes with a tool, I ask a simple question: would I come back?

Not because it's impressive, but because it feels reliable. Because it reduces cognitive load. Because it fades into the background instead of pulling focus.

Many AI products are exciting to try once, but few earn a place in my daily workflow.

Beyond the benchmark

Benchmarks and demos are useful signals, but they rarely explain why one product becomes part of someone's routine while another gets abandoned after a week.

In practice, AI products live or die by how they behave under real conditions: incomplete information, small failures, repeated correction, ongoing constraints. The products that are used long-term are the ones designed with those conditions in mind, not just peak performance.

Evaluating AI tools isn't just about ranking capabilities. It's about whether a product respects how people actually work. That determines whether it gets adopted, trusted, and kept.