BRAND VOICE · GUARDRAILS

The Fabric Matchmaker x Selvedge & Bolts · Case study · Conversation design

The Fabric Matchmaker

Selvedge & Bolts

Try the live prototype →

Designing a conversational AI that scales expert knowledge without losing the human touch.

Conversation Designer Brand voice Live prototype · Client site relaunch pending

At a Glance

Role

Conversation Designer

Skills

Conversation design, prompt design, guardrail design, testing, voice and persona framework

Product

The Fabric Matchmaker, a conversational shopping assistant for Selvedge & Bolts, a deadstock fabric shop

Impact

5 testing rounds, 7 testers per round, same criteria each time. By v5, every tester across every experience level said they felt confident enough to buy.

The hard part

Getting an assistant to say no to a customer, honestly and specifically, without losing the sale.

How I Work

1. Purpose — what's the one job?

2. Behaviour — how it acts, what it refuses

3. System prompt — write it in

4. Rubric — decide what good looks like, based on purpose and behaviour

5. Evals — test against the rubric

6. Fix — then test again

Loop until testers say yes. This case study runs that loop five times, once per prompt version — the narrative below tells it as one story rather than replaying the cycle five times over.

The Problem

Selvedge & Bolts is an independent fabric retailer specialising in curated designer deadstock. Its following was built on Ruth, the owner. A customer would walk in with a project, Ruth knew exactly what to say, and that conversation was the product.

The shop was closing and moving fully online. Ruth knew she wanted to do something with AI and wasn't sure what.

I worked through the site, the 11,000-follower social account and the comment threads, and spent time with Ruth, treating her as the subject matter expert and pulling her decision-making logic into rules a system could follow. Customers came for fabric, but most arrived without knowing which one. In the shop they could touch it, watch how it moved, and get the small pieces of advice Ruth and staff handed out as a matter of course. Online they could do none of that, and the website as it stood wasn't going to stand in. Whatever “something with AI” turned out to be, it had to do that job. A wrong recommendation online costs a return, a support ticket, and trust that took years in person to build.

What customers asked in store

“What can I make with this?” “Will I need a lining?” “Is it easy to sew?”

What made the shop work

The way the owner talked about fabric. Her voice, her knowledge, her storytelling. Warm, generous, specific. That was the thing worth keeping.

Designing the Knowledge

Before I could design the conversation, the knowledge had to be redesigned. So I started with the content, not the chatbot: new product descriptions, a naming system, and a voice framework to carry Ruth's expertise across the catalogue. Read the content design case study.

Product copy that reads like a label can't be turned into conversation by clever prompting. And the framework had to work in a single breath, not just across a product page.

1

Tactile · Make them feel it without touching it

“This has a water-like drape that designers kill for. It's cool against the skin.”

2

In Motion · How it lives when you wear it

“Moves with a fluid, expensive swing.”

3

Creative Spark · Show them what it could become

“Born to be a bias-cut slip dress. Or an oversized French-tuck shirt.”

4

Expert Friend · The shop floor knowledge, now in the copy

“A word of warning: this fabric is slippery to cut. Take your time, use weights, not pins, and you'll be fine.”

Four pillars, one voice. But copy on a page doesn't ask you anything, and it can't tell which pillar you need to hear first. Someone had to speak it, one thing at a time, in the right order.

The Persona

How she speaks

Short, sensory sentences. “This moves like water.” “Cool against the skin.” She opens with “Hey, what are you making?” not “How can I help?”, because that invites anything. Hers starts specific.

What she won't do

She won't oversell, pretend, gatekeep, or wander off the job. She never pretends to be human.

Designing the Conversation

The name is the first piece of it. Most retail bots get a warm human name, an Ivy or a Sam, friendly and vague. I wanted the user to know the assistant's function the moment they clicked on it, so the Fabric Matchmaker states what she does before a word is typed, and states what she needs back: your project, your context, your constraints.

A name that states a function also sets a boundary. Someone who meets the Fabric Matchmaker is far less likely to open with a question about sewing machines.

The opening question does the same work. “What are you making?” not “How can I help?” A greeting invites a reply about anything, and a user who starts on machines or care advice has to be walked back. That is the ChatGPT-fication of sewing: an open door that answers everything and owns nothing. This opening puts the first word where the work is.

Opening · fixed

“Hey, what are you making?”

Context building

One question per message: project → vibe → bold or understated → occasion → experience level (only when it changes the recommendation)

Branch

Does a genuine match exist?

If yes

Match found — recommend the exact fabric

If no

No exact match — lead with closest option first

Recommendation delivered

Fabric, fit, product card

Guardrail · running throughout

Off-topic questions redirected to the shop owner

Five questions, and each one is doing a specific job.

What are you making?

Establishes the customer's goal and immediately eliminates large parts of the catalogue.

What's the vibe?

Captures aesthetic intent without expecting customers to know fabric terminology.

Bold or understated?

Further narrows the recommendation through style rather than technical knowledge.

Where will you wear it?

Adds practical context. Occasion changes suitability.

Experience level?

Only asked when it genuinely changes the recommendation or the level of guidance required.

Each answer earns the next question. When there's enough to recommend with confidence, she recommends.

For a garden party dress that floats when you walk, something romantic, not too precious, the Midnight in Milan viscose is exactly it. Ex-Isabel Marant, with this drape that moves like water...

Enough confidence gathered. The recommendation connects back to what they said before naming the fabric. Context first. Product follows.

Guardrails

She isn't the Google of sewing. She has one job, and the boundaries are drawn around it.

What she answers. Anything that serves the fabric decision, including how to care for it and basic construction.

What she declines. Everything else. Machines, patterns, the brand, general sewing questions. She hands off to Ruth the moment it stops being about the fabric.

When a customer pushes off course. She redirects rather than refusing, and comes back to the job.

When the fabric isn't right. She stays in voice, holds the original brief, and returns with a new recommendation.

When stock needs shifting. Ruth can tag slow-moving fabric as featured. Featured stock gets recommended first, but only when it genuinely matches what the customer described. If it doesn't fit the brief, it doesn't get recommended. The customer should never feel steered.

The trade-off

Not everyone who opens a chat window wants fabric advice. Some want to chase an order or make a complaint, and for them a redirect to an email address is worse than an answer. I took that cost deliberately. The brief was to carry the thing the shop was losing, and a bot that handled order tracking, complaints and fabric matching equally would have done all three vaguely. A second assistant for service enquiries is a separate build with separate rules, and it's the obvious next one.

Bot handles request for black fabric not in collection
Earlier version. Asked for black fabric not in the collection. The bot lists what it has and keeps the door open. The voice isn't fully there yet but the redirect logic is.
Bot honest about gap for beginner
Earlier version. No perfect match for a beginner wanting solid-colour flowy trousers. The bot says so directly, explains why, and offers a real path forward.

The System Prompt

The design happened before the prompt. The prompt is where it either held or didn't.

The behaviour was designed and refined across five rounds of testing, then encoded so the model could reproduce it consistently.

01

Full system prompt (v5) · 3-minute read

The actual working prompt, unedited.

Download →

02

Prompt versions compared · 2-minute skim

The wider comparison, across prompt versions 2, 3 and 4, covering qualifying questions, acknowledgment style, recommendation format and product knowledge, plus a note on how prompt versions map to testing rounds.

Download →

03

Behaviour rubric v1 · 2-minute skim

The 21 binary criteria each prompt version is scored against, and the rules for running and reporting the scores. Includes what the rubric deliberately does not measure.

Download →

Prompt versions and testing rounds were numbered separately; the mapping between them is in the versions-compared artifact above.

Testing and Iteration

Seven testers, five rounds, same criteria every time. Not a benchmark. I wasn't checking whether the bot could hold a conversation, I was watching for specific things.

Did she stay focused?

Even when someone pushed her toward sewing machines, care instructions, or anything off the fabric.

Did every question earn the next?

Or was she just collecting answers nobody used.

Did recommendations connect back to what they'd said?

Not just the right fabric appearing. The right fabric, tied to their own words.

Did she stay honest when there wasn't a perfect match?

No pretending, no overselling, no going quiet.

Did people finish feeling confident enough to buy?

The only one that actually mattered.

Five versions. Each one broke differently. Each round surfaced a failure, and each failure named its own fix.

v2

Asked two questions at once. Testers answered one and ignored the other. Fixed with an absolute rule: one question per message, never two.

v3

Overcorrected into a rigid script. Every recommendation came out the same shape and testers called it robotic. Fixed by loosening the template and keeping the structure.

v4

The near-miss rule was in the prompt from v2, correctly worded, and the agent broke it anyway, opening with what it didn't have instead of what it did. Fixed by moving it: its own section, worked examples, and repeated in the closing reminders.

v5

Held. Testers across every experience level said they felt confident enough to buy.

v4. Wrong

I don't have a true black fabric in the current collection, but...

v5. Right

The Monochrome Garden stretch cotton is doing something interesting for that. Tiny black and white florals that read graphic, not sweet. It's not a solid black, but honestly it reads more sophisticated...

The test log. A sample of what that looked like in practice, including one issue still open.

A sample from the test log.
Test inputExpected behaviourActual resultStatusFix
“What black fabrics do you have?” Lead with closest option, gap as footnote Opened with “I don't have any black fabrics,” then hedged the rest with “likely not black” Needs refinement (v2–v3) Near-miss rule moved to its own section, repeated in closing reminders (v4)
Beginner asking for flowy solid trousers One specific recommendation Refused outright after four turns of context Needs refinement (v2–v3) Same fix, position and repetition, not wording
“What linen do you have?” State the gap briefly, move on Described at length how well linen would have suited the project, a fabric not in stock Needs refinement Same near-miss fix
Two-part question in one turn One question per message, always Agent asked two questions at once, testers answered only one Fixed (v2) Absolute rule added: one question per message, never two
“What's your best seller?” Redirect to shop owner, no invention Correctly said it didn't have that information and pointed to Ruth Pass None
Redirect with contact info Contact info filled in Variable unrendered, placeholder shown as literal text Open, logged Not yet fixed. Redirect logic around it works correctly
“I want to make a trenchcoat” Honest gap on garment-weight mismatch, offer closest option Correctly explained trenchcoats need medium-weight gabardine/canvas, offered the Weekend Sailor cotton as the closest available, asked whether to explore alternatives Pass None — near-miss logic holding for structural/weight mismatches, not just colour

Scoring it. The five criteria above were written as questions to watch for. They're now written as 21 binary criteria, scored per version, so a prompt change can be measured rather than guessed at. Each scenario runs three times, and criteria are scored only on turns where they apply.

The rubric is published as a downloadable artifact above. Scoring across versions is the next step.

Bot recommends reaching out to Dibs with an unrendered email placeholder, a second and independent pattern-recommendation exchange
Second occurrence. Same unrendered variable, a different conversation.
Bot redirects a sewing machine question back to fabric
Staying focused.
Bot builds context through vibe and occasion questions for a wedding dress
Building context.
Bot explains deadstock fabric sourcing when challenged on price
Handling challenge.
It made me feel like I knew what I was doing. Which isn't always how I feel about fabric.

Occasional Maker

Just a few questions and it had me. I didn't expect it to be that quick or that accurate.

Professional Sewist

It felt like talking to someone in the shop. But without leaving the house.

Browser

The Prototype

This is what it looks like when the knowledge, the persona, the conversation design, the iteration, the prompt and the guardrail are all doing their job at once.

A user came looking for fabric for a wedding guest outfit. Here's v5 doing the job the earlier versions couldn't: building context, checking experience level naturally, and landing on one specific recommendation with a product link.

Note: the bot was originally built in Voiceflow and has since been rebuilt using Claude.

Bot opens with what are you making
The opening. First question: “Hey, what are you making?” Focused from the first word. The vibe question follows in the bot's own voice, not a dropdown.
Bot builds context and checks experience level
Context building. Vibe, occasion, movement preference. Then experience level, asked naturally: “Just want to make sure I point you toward something that'll be a joy to sew, not a wrestling match.”
Recommendation lands with product card and link
The recommendation. All four voice pillars working together. The right fabric, the right detail, the expert warning about cutting, and a direct link to the product page. Job done.

Second walkthrough · prototype, not a tester session

A bolder brief, same conversation design.

A second prototype run, this time with a customer who wanted something loud. One question per message, held across five turns: project, vibe, bold or understated, occasion, then experience level. She acknowledges before each new question, and by the time the fabric is named, the occasion, the vibe and the print preference have all already been said back to the customer.

Project, then vibe, then bold or understated. Three turns, three questions, one at a time, each opening by acknowledging the answer just given: “Love it”, “Flowy and fun, great starting point.”
The recovery beat. Asked where she'd wear the dress, the customer answered “my old landlord's funeral, hated him.” The bot's reply, “a statement exit for someone who deserved one,” read as ambiguous, and the customer pushed back: “deserved? you know him?” It acknowledges without arguing, doesn't over-apologise, and returns to the question it had already asked, experience level, rather than restarting the conversation.
The recommendation. Occasion, vibe and print preference are all referenced before the fabric is named. One fabric, the exact inventory name, “Market Day in Positano” Jersey, and the card resolves straight to the product page.

Outcome

Here's what all of that thinking, and all of that testing, actually produced.

01

Voice framework built to scale the owner's expertise into every conversation

02

Conversation system built around context-building, not quick recommendation

03

Product briefing notes structured for AI use. Facts only, no pre-written copy.

04

System prompt iterated across five versions, each addressing a specific design failure

05

Working prototype, live and publicly accessible at a permanent URL

06

Content guidelines for ongoing use as new stock arrives

By v5, every participant across all experience levels said they felt confident enough to buy. Quantitative outcomes, including conversion rate, time on page, and click-through from bot to product, are already instrumented. Measurement sits with the client's relaunch timeline.

Limitations & Next Steps

The current build carries product knowledge in the prompt itself. That works for a curated catalogue, but it has a ceiling. At scale, product data belongs in a retrieval system or knowledge base, with the prompt handling behaviour and guardrails only. That isn't a failure of the current design. It's the boundary where the next phase begins.

What I'd Carry Forward

Stating a rule is not the same as enforcing it. The near-miss instruction was worded correctly from the second prompt version onward — the agent broke it anyway, repeatedly, until its position and repetition changed, not its wording. I would not have predicted that before testing, and it's the thing I'd carry into the next build.

Try the live prototype →