TokenMind

Don't wanna read? Listen my story!!

TokenMind is a real-time prompt quality analyzer built directly inside Wibey, Walmart's AI assistant used by 10,000+ engineers every day. It scores prompts 1–100 as you type, flags up to 7 types of issues, and routes you to the right AI model — all before you hit send. It was designed and prototyped in one week during a Walmart internal hackathon.

Role:

UX Designer II

Duration:

1 week

The Challenge

01

How might we help engineers write better AI prompts in the moment they're typing, so they get higher-quality responses without needing to learn prompt engineering first?

Every day, 10,000+ engineers at Walmart open Wibey and type a prompt. Most of them are doing it wrong without knowing it.


Not because they don't care. Because there is no feedback. A prompt goes in. An output comes back. If the output is vague, incomplete, or off-target, the engineer tries again. And again. Each retry burns tokens. Each poorly scoped prompt gets routed to a model that wasn't built for the task. And because the feedback only arrives after the output, the cost is already spent by the time anyone notices something went wrong.

Prompt engineering is a skill. Most engineers don't have it, and you can't train 10,000 people to have it fast enough to matter. The gap between what engineers type and what the model needs to perform well is real, consistent, and expensive.


The problem, precisely:

  • Engineers write prompts with vague language, missing context, no format, no persona, no constraints, no examples, or over-length — often without realizing any of it

  • There is no real-time feedback mechanism between typing and sending

  • The only signal is the output — and by then, the tokens are spent

  • Bad prompts also get routed to the wrong models, compounding the waste

  • Estimated cost: $100K–$500K per year in wasted tokens and model overspend across Walmart engineering

The scale made the silence expensive.

Solutions

02

TokenMind

TokenMind lives where the prompt is written. No new tool, no new tab, no training course. It is embedded directly inside Wibey the tool engineers are already using and it works in real time, as the engineer types.


How it works:

Prompt Scoring: Every prompt is rated 1–100 the moment it is typed. Not after submission. Not after a bad output. In the moment, before send.


7 Issue Types Flagged: TokenMind detects and surfaces seven categories of prompt problems — vague language, missing context, no format specified, no persona, no constraints, no examples, and over-length. Each issue is identified so the engineer knows exactly what to fix, not just that something is wrong.


Right Model, First Time: Based on what the prompt is asking for, TokenMind routes it to the AI model best suited for the task. No guessing. No trial and error.


Zero expertise required: The intelligence is in the interface. An engineer who has never heard the words "prompt engineering" gets the same quality feedback as one who has. TokenMind scales because it does not depend on the user already knowing what good looks like.

From vague to optimized in under 10 seconds.

Impact

03

Why this matters.

Annual savings potential

~$500K

Prompt issue types detected

7

Prompt retries

Zero — fix before you send

Expertise required

None

Engineers reached

10,000+

Reflection

04

What I’d take forward.

One week forces decisions. There is no time to refine indefinitely you ship the clearest version of the idea, not the most polished one.


The design decision that held the project together was the decision to embed inside Wibey rather than build something standalone. A separate tool would have required behavior change. Inside Wibey, it required none. The engineer is already there.


The prompt is already being typed. TokenMind just makes that moment smarter.

The insight was not technical. It was about where the feedback needs to live. Not in a dashboard, not in a training deck, not in a course — in the input field, at the exact moment the engineer is about to make a mistake.

TokenMind

Don't wanna read? Listen my story!!

TokenMind is a real-time prompt quality analyzer built directly inside Wibey, Walmart's AI assistant used by 10,000+ engineers every day. It scores prompts 1–100 as you type, flags up to 7 types of issues, and routes you to the right AI model — all before you hit send. It was designed and prototyped in one week during a Walmart internal hackathon.

Role:

UX Designer II

Duration:

1 week

The Challenge

01

How might we help engineers write better AI prompts in the moment they're typing, so they get higher-quality responses without needing to learn prompt engineering first?

Every day, 10,000+ engineers at Walmart open Wibey and type a prompt. Most of them are doing it wrong without knowing it.


Not because they don't care. Because there is no feedback. A prompt goes in. An output comes back. If the output is vague, incomplete, or off-target, the engineer tries again. And again. Each retry burns tokens. Each poorly scoped prompt gets routed to a model that wasn't built for the task. And because the feedback only arrives after the output, the cost is already spent by the time anyone notices something went wrong.

Prompt engineering is a skill. Most engineers don't have it, and you can't train 10,000 people to have it fast enough to matter. The gap between what engineers type and what the model needs to perform well is real, consistent, and expensive.


The problem, precisely:

  • Engineers write prompts with vague language, missing context, no format, no persona, no constraints, no examples, or over-length — often without realizing any of it

  • There is no real-time feedback mechanism between typing and sending

  • The only signal is the output — and by then, the tokens are spent

  • Bad prompts also get routed to the wrong models, compounding the waste

  • Estimated cost: $100K–$500K per year in wasted tokens and model overspend across Walmart engineering

The scale made the silence expensive.

Solutions

02

TokenMind

TokenMind lives where the prompt is written. No new tool, no new tab, no training course. It is embedded directly inside Wibey the tool engineers are already using and it works in real time, as the engineer types.


How it works:

Prompt Scoring: Every prompt is rated 1–100 the moment it is typed. Not after submission. Not after a bad output. In the moment, before send.


7 Issue Types Flagged: TokenMind detects and surfaces seven categories of prompt problems — vague language, missing context, no format specified, no persona, no constraints, no examples, and over-length. Each issue is identified so the engineer knows exactly what to fix, not just that something is wrong.


Right Model, First Time: Based on what the prompt is asking for, TokenMind routes it to the AI model best suited for the task. No guessing. No trial and error.


Zero expertise required: The intelligence is in the interface. An engineer who has never heard the words "prompt engineering" gets the same quality feedback as one who has. TokenMind scales because it does not depend on the user already knowing what good looks like.

From vague to optimized in under 10 seconds.

Impact

03

Why this matters.

Metric

Before

Annual savings potential

~$500K

Prompt issue types

detected

7

Prompt retries

Zero — fix before you send

Expertise required

None

Engineers reached

10,000+

Reflection

04

What I’d take forward.

One week forces decisions. There is no time to refine indefinitely you ship the clearest version of the idea, not the most polished one.


The design decision that held the project together was the decision to embed inside Wibey rather than build something standalone. A separate tool would have required behavior change. Inside Wibey, it required none. The engineer is already there.


The prompt is already being typed. TokenMind just makes that moment smarter.

The insight was not technical. It was about where the feedback needs to live. Not in a dashboard, not in a training deck, not in a course — in the input field, at the exact moment the engineer is about to make a mistake.

TokenMind

Don't wanna read? Listen my story!!

TokenMind is a real-time prompt quality analyzer built directly inside Wibey, Walmart's AI assistant used by 10,000+ engineers every day. It scores prompts 1–100 as you type, flags up to 7 types of issues, and routes you to the right AI model — all before you hit send. It was designed and prototyped in one week during a Walmart internal hackathon.

Role:

UX Designer II

Duration:

1 week

The Challenge

01

How might we help engineers write better AI prompts in the moment they're typing, so they get higher-quality responses without needing to learn prompt engineering first?

Every day, 10,000+ engineers at Walmart open Wibey and type a prompt. Most of them are doing it wrong without knowing it.


Not because they don't care. Because there is no feedback. A prompt goes in. An output comes back. If the output is vague, incomplete, or off-target, the engineer tries again. And again. Each retry burns tokens. Each poorly scoped prompt gets routed to a model that wasn't built for the task. And because the feedback only arrives after the output, the cost is already spent by the time anyone notices something went wrong.

Prompt engineering is a skill. Most engineers don't have it, and you can't train 10,000 people to have it fast enough to matter. The gap between what engineers type and what the model needs to perform well is real, consistent, and expensive.


The problem, precisely:

  • Engineers write prompts with vague language, missing context, no format, no persona, no constraints, no examples, or over-length — often without realizing any of it

  • There is no real-time feedback mechanism between typing and sending

  • The only signal is the output — and by then, the tokens are spent

  • Bad prompts also get routed to the wrong models, compounding the waste

  • Estimated cost: $100K–$500K per year in wasted tokens and model overspend across Walmart engineering

The scale made the silence expensive.

Solutions

02

TokenMind

TokenMind lives where the prompt is written. No new tool, no new tab, no training course. It is embedded directly inside Wibey the tool engineers are already using and it works in real time, as the engineer types.


How it works:

Prompt Scoring: Every prompt is rated 1–100 the moment it is typed. Not after submission. Not after a bad output. In the moment, before send.


7 Issue Types Flagged: TokenMind detects and surfaces seven categories of prompt problems — vague language, missing context, no format specified, no persona, no constraints, no examples, and over-length. Each issue is identified so the engineer knows exactly what to fix, not just that something is wrong.


Right Model, First Time: Based on what the prompt is asking for, TokenMind routes it to the AI model best suited for the task. No guessing. No trial and error.


Zero expertise required: The intelligence is in the interface. An engineer who has never heard the words "prompt engineering" gets the same quality feedback as one who has. TokenMind scales because it does not depend on the user already knowing what good looks like.

From vague to optimized in under 10 seconds.

Impact

03

Why this matters.

Metric

Value

Annual savings potential

~$500K

Prompt issue types detected

7

Prompt retries

Zero — fix before you send

Expertise required

None

Engineers reached

10,000+

Reflection

04

What I’d take forward.

One week forces decisions. There is no time to refine indefinitely you ship the clearest version of the idea, not the most polished one.


The design decision that held the project together was the decision to embed inside Wibey rather than build something standalone. A separate tool would have required behavior change. Inside Wibey, it required none. The engineer is already there.


The prompt is already being typed. TokenMind just makes that moment smarter.

The insight was not technical. It was about where the feedback needs to live. Not in a dashboard, not in a training deck, not in a course — in the input field, at the exact moment the engineer is about to make a mistake.