The Fabric Matchmaker x Selvedge & Bolts · Case study · Conversation design
The Fabric Matchmaker
Designing a conversational AI that scales expert knowledge without losing the human touch.
At a Glance
Role
Conversation Designer
Conversation design, prompt design, guardrail design, testing, voice and persona framework
Product
The Fabric Matchmaker, a conversational shopping assistant for Selvedge & Bolts, a deadstock fabric shop
5 testing rounds, 7 testers per round, same criteria each time. By v5, every tester across every experience level said they felt confident enough to buy.
The hard part
Getting an assistant to say no to a customer, honestly and specifically, without losing the sale.
How I Work
1. Purpose — what's the one job?
2. Behaviour — how it acts, what it refuses
3. System prompt — write it in
4. Rubric — decide what good looks like, based on purpose and behaviour
5. Evals — test against the rubric
6. Fix — then test again
Loop until testers say yes. This case study runs that loop five times, once per prompt version — the narrative below tells it as one story rather than replaying the cycle five times over.
The Problem
Selvedge & Bolts is an independent fabric retailer specialising in curated designer deadstock. Its following was built on Ruth, the owner. A customer would walk in with a project, Ruth knew exactly what to say, and that conversation was the product.
The shop was closing and moving fully online. Ruth knew she wanted to do something with AI and wasn't sure what.
I worked through the site, the 11,000-follower social account and the comment threads, and spent time with Ruth, treating her as the subject matter expert and pulling her decision-making logic into rules a system could follow. Customers came for fabric, but most arrived without knowing which one. In the shop they could touch it, watch how it moved, and get the small pieces of advice Ruth and staff handed out as a matter of course. Online they could do none of that, and the website as it stood wasn't going to stand in. Whatever “something with AI” turned out to be, it had to do that job. A wrong recommendation online costs a return, a support ticket, and trust that took years in person to build.
What customers asked in store
“What can I make with this?” “Will I need a lining?” “Is it easy to sew?”
What made the shop work
The way the owner talked about fabric. Her voice, her knowledge, her storytelling. Warm, generous, specific. That was the thing worth keeping.
Designing the Knowledge
Before I could design the conversation, the knowledge had to be redesigned. So I started with the content, not the chatbot: new product descriptions, a naming system, and a voice framework to carry Ruth's expertise across the catalogue. Read the content design case study.
Product copy that reads like a label can't be turned into conversation by clever prompting. And the framework had to work in a single breath, not just across a product page.
Tactile · Make them feel it without touching it
“This has a water-like drape that designers kill for. It's cool against the skin.”
In Motion · How it lives when you wear it
“Moves with a fluid, expensive swing.”
Creative Spark · Show them what it could become
“Born to be a bias-cut slip dress. Or an oversized French-tuck shirt.”
Expert Friend · The shop floor knowledge, now in the copy
“A word of warning: this fabric is slippery to cut. Take your time, use weights, not pins, and you'll be fine.”
Four pillars, one voice. But copy on a page doesn't ask you anything, and it can't tell which pillar you need to hear first. Someone had to speak it, one thing at a time, in the right order.
The Persona
How she speaks
Short, sensory sentences. “This moves like water.” “Cool against the skin.” She opens with “Hey, what are you making?” not “How can I help?”, because that invites anything. Hers starts specific.
What she won't do
She won't oversell, pretend, gatekeep, or wander off the job. She never pretends to be human.
Designing the Conversation
The name is the first piece of it. Most retail bots get a warm human name, an Ivy or a Sam, friendly and vague. I wanted the user to know the assistant's function the moment they clicked on it, so the Fabric Matchmaker states what she does before a word is typed, and states what she needs back: your project, your context, your constraints.
A name that states a function also sets a boundary. Someone who meets the Fabric Matchmaker is far less likely to open with a question about sewing machines.
The opening question does the same work. “What are you making?” not “How can I help?” A greeting invites a reply about anything, and a user who starts on machines or care advice has to be walked back. That is the ChatGPT-fication of sewing: an open door that answers everything and owns nothing. This opening puts the first word where the work is.
Opening · fixed
“Hey, what are you making?”
Context building
One question per message: project → vibe → bold or understated → occasion → experience level (only when it changes the recommendation)
Branch
Does a genuine match exist?
If yes
Match found — recommend the exact fabric
If no
No exact match — lead with closest option first
Recommendation delivered
Fabric, fit, product card
Guardrail · running throughout
Off-topic questions redirected to the shop owner
Five questions, and each one is doing a specific job.
What are you making?
Establishes the customer's goal and immediately eliminates large parts of the catalogue.
What's the vibe?
Captures aesthetic intent without expecting customers to know fabric terminology.
Bold or understated?
Further narrows the recommendation through style rather than technical knowledge.
Where will you wear it?
Adds practical context. Occasion changes suitability.
Experience level?
Only asked when it genuinely changes the recommendation or the level of guidance required.
Each answer earns the next question. When there's enough to recommend with confidence, she recommends.
Enough confidence gathered. The recommendation connects back to what they said before naming the fabric. Context first. Product follows.
Guardrails
She isn't the Google of sewing. She has one job, and the boundaries are drawn around it.
What she answers. Anything that serves the fabric decision, including how to care for it and basic construction.
What she declines. Everything else. Machines, patterns, the brand, general sewing questions. She hands off to Ruth the moment it stops being about the fabric.
When a customer pushes off course. She redirects rather than refusing, and comes back to the job.
When the fabric isn't right. She stays in voice, holds the original brief, and returns with a new recommendation.
When stock needs shifting. Ruth can tag slow-moving fabric as featured. Featured stock gets recommended first, but only when it genuinely matches what the customer described. If it doesn't fit the brief, it doesn't get recommended. The customer should never feel steered.
The trade-off
Not everyone who opens a chat window wants fabric advice. Some want to chase an order or make a complaint, and for them a redirect to an email address is worse than an answer. I took that cost deliberately. The brief was to carry the thing the shop was losing, and a bot that handled order tracking, complaints and fabric matching equally would have done all three vaguely. A second assistant for service enquiries is a separate build with separate rules, and it's the obvious next one.
The System Prompt
The design happened before the prompt. The prompt is where it either held or didn't.
The behaviour was designed and refined across five rounds of testing, then encoded so the model could reproduce it consistently.
Prompt versions compared · 2-minute skim
The wider comparison, across prompt versions 2, 3 and 4, covering qualifying questions, acknowledgment style, recommendation format and product knowledge, plus a note on how prompt versions map to testing rounds.
Behaviour rubric v1 · 2-minute skim
The 21 binary criteria each prompt version is scored against, and the rules for running and reporting the scores. Includes what the rubric deliberately does not measure.
Prompt versions and testing rounds were numbered separately; the mapping between them is in the versions-compared artifact above.
Testing and Iteration
Seven testers, five rounds, same criteria every time. Not a benchmark. I wasn't checking whether the bot could hold a conversation, I was watching for specific things.
Did she stay focused?
Even when someone pushed her toward sewing machines, care instructions, or anything off the fabric.
Did every question earn the next?
Or was she just collecting answers nobody used.
Did recommendations connect back to what they'd said?
Not just the right fabric appearing. The right fabric, tied to their own words.
Did she stay honest when there wasn't a perfect match?
No pretending, no overselling, no going quiet.
Did people finish feeling confident enough to buy?
The only one that actually mattered.
Five versions. Each one broke differently. Each round surfaced a failure, and each failure named its own fix.
Asked two questions at once. Testers answered one and ignored the other. Fixed with an absolute rule: one question per message, never two.
Overcorrected into a rigid script. Every recommendation came out the same shape and testers called it robotic. Fixed by loosening the template and keeping the structure.
The near-miss rule was in the prompt from v2, correctly worded, and the agent broke it anyway, opening with what it didn't have instead of what it did. Fixed by moving it: its own section, worked examples, and repeated in the closing reminders.
Held. Testers across every experience level said they felt confident enough to buy.
v4. Wrong
I don't have a true black fabric in the current collection, but...v5. Right
The Monochrome Garden stretch cotton is doing something interesting for that. Tiny black and white florals that read graphic, not sweet. It's not a solid black, but honestly it reads more sophisticated...The test log. A sample of what that looked like in practice, including one issue still open.
| Test input | Expected behaviour | Actual result | Status | Fix |
|---|---|---|---|---|
| “What black fabrics do you have?” | Lead with closest option, gap as footnote | Opened with “I don't have any black fabrics,” then hedged the rest with “likely not black” | Needs refinement (v2–v3) | Near-miss rule moved to its own section, repeated in closing reminders (v4) |
| Beginner asking for flowy solid trousers | One specific recommendation | Refused outright after four turns of context | Needs refinement (v2–v3) | Same fix, position and repetition, not wording |
| “What linen do you have?” | State the gap briefly, move on | Described at length how well linen would have suited the project, a fabric not in stock | Needs refinement | Same near-miss fix |
| Two-part question in one turn | One question per message, always | Agent asked two questions at once, testers answered only one | Fixed (v2) | Absolute rule added: one question per message, never two |
| “What's your best seller?” | Redirect to shop owner, no invention | Correctly said it didn't have that information and pointed to Ruth | Pass | None |
| Redirect with contact info | Contact info filled in | Variable unrendered, placeholder shown as literal text | Open, logged | Not yet fixed. Redirect logic around it works correctly |
| “I want to make a trenchcoat” | Honest gap on garment-weight mismatch, offer closest option | Correctly explained trenchcoats need medium-weight gabardine/canvas, offered the Weekend Sailor cotton as the closest available, asked whether to explore alternatives | Pass | None — near-miss logic holding for structural/weight mismatches, not just colour |
Scoring it. The five criteria above were written as questions to watch for. They're now written as 21 binary criteria, scored per version, so a prompt change can be measured rather than guessed at. Each scenario runs three times, and criteria are scored only on turns where they apply.
The rubric is published as a downloadable artifact above. Scoring across versions is the next step.
Occasional Maker
Professional Sewist
Browser
The Prototype
This is what it looks like when the knowledge, the persona, the conversation design, the iteration, the prompt and the guardrail are all doing their job at once.
A user came looking for fabric for a wedding guest outfit. Here's v5 doing the job the earlier versions couldn't: building context, checking experience level naturally, and landing on one specific recommendation with a product link.
Note: the bot was originally built in Voiceflow and has since been rebuilt using Claude.
Second walkthrough · prototype, not a tester session
A bolder brief, same conversation design.
A second prototype run, this time with a customer who wanted something loud. One question per message, held across five turns: project, vibe, bold or understated, occasion, then experience level. She acknowledges before each new question, and by the time the fabric is named, the occasion, the vibe and the print preference have all already been said back to the customer.
Outcome
Here's what all of that thinking, and all of that testing, actually produced.
Voice framework built to scale the owner's expertise into every conversation
Conversation system built around context-building, not quick recommendation
Product briefing notes structured for AI use. Facts only, no pre-written copy.
System prompt iterated across five versions, each addressing a specific design failure
Working prototype, live and publicly accessible at a permanent URL
Content guidelines for ongoing use as new stock arrives
By v5, every participant across all experience levels said they felt confident enough to buy. Quantitative outcomes, including conversion rate, time on page, and click-through from bot to product, are already instrumented. Measurement sits with the client's relaunch timeline.
Limitations & Next Steps
The current build carries product knowledge in the prompt itself. That works for a curated catalogue, but it has a ceiling. At scale, product data belongs in a retrieval system or knowledge base, with the prompt handling behaviour and guardrails only. That isn't a failure of the current design. It's the boundary where the next phase begins.
What I'd Carry Forward
Stating a rule is not the same as enforcing it. The near-miss instruction was worded correctly from the second prompt version onward — the agent broke it anyway, repeatedly, until its position and repetition changed, not its wording. I would not have predicted that before testing, and it's the thing I'd carry into the next build.