A quoting pilot should tell you whether your reps can finish real requests with less work, including the time they spend checking and correcting AI drafts. A fast demonstration on a clean parts list does not answer that question.
The scorecard below compares manual preparation with the complete assisted process. Start with a representative set of RFQs, keep some out of setup, and record what happened to every request, including the ones that needed clarification or manual work.
1. Select representative RFQs for the pilot
Start with 20–30 RFQs across your common formats and difficult cases. Add more examples wherever that set misses a request type or a costly exception.
Group requests by both document-reading difficulty and commercial complexity. An easy scan can contain a difficult commercial decision, while a messy scan may contain exact catalog numbers.
- Clean requests with exact catalog numbers.
- Customer aliases, old part numbers, and “same as last time” requests.
- Different units, pack sizes, or quantities within the same file.
- Expired discounts and customer-specific terms.
- Missing information, conflicting attachments, and substitutions that need review.
2. Separate setup examples from evaluation examples
Provide some RFQs and their completed quotes for configuration. Reserve others for evaluation. If every test request appeared during setup, you have not measured how the agent handles a new request.
Record the files available to both the rep and the agent. A comparison is difficult to interpret if one has the current price list and the other only has an old quote.
3. Agree pass criteria before seeing results
Define material errors with the commercial owner: wrong customer, specification, unit or unauthorized terms. Set a maximum acceptable review time for each request type. State when the correct result is a clarification rather than a completed quote.
Use the consequences in your plant to distinguish errors. ASQ’s cost-of-quality model separates checking and prevention from failures such as scrap, rework and warranty claims. A misplaced address line and a wrong material specification should not carry the same weight in the scorecard.
For a machined component, include a request where the buyer changes the drawing revision without changing the item number. For a distributor, include a case quantity change with the same customer description. GS1’s pack-quantity rules explain why package identity can change when case contents change. Score whether the workflow detects the changed requirement and asks for the right decision, not just whether it produces a plausible quote.
Keep every evaluation attempt in the results, including failures and manual fallbacks. Report counts with denominators. If 18 of 20 requests reach approval, explain what happened to the other two rather than reporting only the 18 completed cases.
4. Measure quote accuracy and review time
| Measure | How to record it |
|---|---|
| Manual preparation time | Active minutes spent preparing the quote, including lookups. |
| Agent draft review time | Active minutes checking and correcting the draft, including escalations. |
| Material errors | Wrong customer, part, quantity, unit, price, or commercial condition. |
| Unresolved items | Decisions correctly left open and the information needed to resolve them. |
| Approved result | Whether the quote was ready to send after review. |
5. Test how the agent applies reviewer feedback
Correct one customer-specific decision. Use a new request from that customer to test reuse. Then use a similar request from another customer to test whether the agent applies the exception too broadly. Record both outcomes.
For each correction, retain the customer, source and scope of the decision. A one-time project concession must not become the default for every later quote.
For historical pricing, test whether the old quote helps identify the part while the current price still receives the right check. Remembering an old value is different from knowing that it remains valid.
6. Decide whether the pilot supports a live deployment
Review total time, error types, unresolved work, and how much setup was required. Agree which decisions need human approval in production. Verify the native connections, permissions, operating ownership, and cost used in the pilot before expanding volume or authority.
Test the configured integrations separately from draft quality. Confirm that an approved quote reaches the correct ERP record and that a retry does not create duplicates. Name the owner who will review failures after launch.
Turn results into a rollout decision
For each request group, report manual minutes, review minutes, material errors and unresolved decisions. Include setup effort and recurring costs in the rollout budget, including the internal work the vendor does not charge for. Keep released capacity separate from payroll savings.
Require a live transaction test before expanding authority. For example, the Business Central quote API represents a specific business record; generating a correct PDF alone does not establish that the integration can create it.
Run your first quoting project with Bourne
You do not need to build the evaluation and the system alone. We work with your team to configure the workflow, connect the required systems through Bourne's native integrations and test the result on your requests. Bring completed quotes and the cases that still take your senior reps too long. Plan your Bourne deployment with a clear first workflow and a scorecard everyone can use.
Frequently asked questions
Are 20 RFQs enough to prove production accuracy?
No. A small sample can expose problems and help scope the project, but it cannot establish a reliable error rate for rare cases. Expand the evaluation and monitor live results before increasing autonomy.
Should the vendor see all our test answers during setup?
No. Provide examples for configuration and reserve a separate set for evaluation. If the test only repeats examples used during setup, it does not measure how the workflow handles new work.
Is asking for human review a pilot failure?
Not when the request genuinely needs approval or missing information. Record correct escalations separately from avoidable ones, and include the human time in the result.
What should make us stop a rollout?
Stop expansion when the workflow repeatedly makes material errors, bypasses approval or cannot recover predictably from failures. Narrow the scope and resolve the cause before retesting on fresh requests.
Further reading
ASQ: cost of quality
Separating prevention and appraisal costs from the costs of failures.
GS1: pack and case quantity changes
When changed case or pallet quantities require a new GTIN.
The Manufacturing AI Roadmap
Decide what to automate next.
Use our free guide to assess your commercial operations and choose a starting point for AI.