Businesses are investing heavily in AI, but many still struggle to turn that investment into measurable value.

Read the full whitepaper here

The problem is usually not the technology. It’s that teams start with a model, feature, or executive mandate instead of a clearly defined human need. They ask, “What can we build with AI?” when they should ask, “What outcome do people need, and can AI help us deliver it?”

A practical way to answer that question is to combine three ideas:

  1. Define the job the AI is being hired to do.
  2. Design the experience around the people who will use it.
  3. Establish how you’ll evaluate success before development begins.

Start with the job, not the technology

The Jobs to Be Done framework focuses on the outcome someone is trying to achieve. People don’t buy a drill because they want a drill. They buy it because they need a hole.

The same principle applies to AI.

“We need an AI chatbot” isn’t a clearly defined job. The actual job might be to provide customers with immediate answers, reduce support volume by 30%, or help employees find accurate policy information without contacting HR.

Framing the project around that outcome shifts the conversation from what AI can do to what the business and its users actually need.

The human experience determines the value

Identifying the right job is only the beginning. The AI system must also be usable, trustworthy, and adopted by the people it’s meant to help.

Users don’t care how sophisticated the underlying model is. They care whether the system makes their work easier and produces results they can trust.

That means involving users early, understanding their context, testing prototypes, and improving the experience based on feedback. It also means giving users appropriate transparency, control, and fallback options.

An AI system can be technically impressive and still produce no business value if people don’t use it.

Define the “Outcomes to Be Eval’d”

Once the job is clear, the next question is:

How will we know whether the AI is doing that job successfully?

This is what I call “Outcomes to Be Eval’d.”

Traditional software is generally deterministic. AI is probabilistic. Its answers can vary, edge cases can appear unexpectedly, and a system that performs well in a demo may struggle in the real world.

Evaluation therefore can’t be something added at the end. It needs to shape development from the beginning.

The process is straightforward:

  1. Define success measures. Identify the quantitative and qualitative outcomes that matter.
  2. Establish evaluation criteria. Decide how each outcome will be tested and what constitutes a passing result.
  3. Develop with evaluation feedback. Build the smallest useful version and test it immediately.
  4. Continuously refine. Monitor performance after deployment and adapt as new situations emerge.

Consider an AI-powered HR assistant that answers questions about company holidays. Its job is not simply to generate a plausible response. Its job is to quickly provide the correct information from the official employee handbook.

That job could be evaluated using:

  • A keyword test confirming that required holidays are included
  • An AI judge assessing completeness and accuracy
  • A latency test confirming that the response arrives within one second

The assistant might pass the accuracy test but fail the latency requirement. That distinction matters. Without explicit evaluations, the team might see a correct answer and declare the project successful, even though it failed an important part of the user experience.

Moving from AI hype to business impact

Successful AI projects begin with empathy for the person using the system and end with evidence that the intended outcome was achieved.

Jobs to Be Done gives the project a clear purpose. Human-centered design makes the solution usable and trustworthy. Outcomes to Be Eval’d connects that purpose to repeatable, measurable performance.

Together, these ideas create a common language between business stakeholders, designers, and technical teams. They move the conversation away from models and features and toward the question that ultimately matters:

Did the AI help someone achieve a valuable outcome, and can we prove it?