Skip to main content

PF_MULTI_005 — Multilingual quote batch accuracy and start-to-send timing

This proof tests one bounded 1,000-request multilingual quote batch. It requires one independently reviewable draft per source, no more than 50 mismatched complete quotes before correction, native draft storage for accepted submissions, and an aggregate start-to-send result no greater than one-third of the comparable prior-process aggregate. It does not establish universal language coverage, correctness without review, or capacity beyond the tested batch.

kind: proof_catalog
schema_version: 2
proofs:
- id: PF_MULTI_005
vertical_ids:
- M1_V1
- M1_V2
- M1_V3
- M1_V4
- M1_V5
- M1_V6
- M1_V7
name: 'Multilingual quote batch accuracy and start-to-send timing'
status: approved
proof_type: 'dataset test'
safe_artifact: >-
One sanitized, permissioned evaluation package containing exactly 1,000
representative multilingual quote-request inputs, one source identifier
per request, the required quote-field list, expected values derived from
the reported real-database truth set, and start and send timestamps. The
package excludes customer-identifying and unnecessary commercial data;
no live customer database access is required.
procedure:
- >-
Record the represented languages, more-than-25-production-customer
scope, required quote fields, source identifiers, and expected values
for all 1,000 requests before processing.
- >-
Declare the prior-process comparison set, start and send timestamp
definitions, and one aggregation method that will be applied unchanged
to both the prior and configured-workflow sets.
- >-
Upload the 1,000 requests through the configured multilingual quote
batch intake.
- >-
Inspect whether the intake returns exactly one source-associated,
independently reviewable structured quote draft for every request.
- >-
Before manual correction, compare every required field in each draft
with its expected real-database truth-set value. Mark the complete quote
mismatched when any required field differs or is missing.
- >-
Count matched and mismatched complete quotes, preserving the mismatch
reason for reviewer inspection.
- >-
Have a reviewer confirm or correct drafts. Submit reviewer-accepted
valid drafts through the native quote creation path and inspect the
resulting stored sales/quote drafts.
- >-
Record the configured workflow's send timestamp for each sent quote and
compare its aggregate start-to-send elapsed time with the prior-process
aggregate using the same predeclared method.
expected_observation: >-
The evaluation exposes one independently reviewable structured draft per
source request, a complete-quote match result against the expected
real-database values, stored sales/quote drafts for reviewer-accepted
submissions, and comparable aggregate start-to-send timing for the prior
and configured workflows.
acceptance_criterion: >-
The 1,000-request batch returns exactly 1,000 source-associated,
independently reviewable drafts; no more than 50 complete quotes are
mismatched before manual correction, where any missing or differing
required field fails the complete quote; every reviewer-accepted valid
draft selected for submission produces an inspectable stored sales/quote
draft; and, under this proof's explicit interpretation of "3x faster,"
the configured workflow's aggregate start-to-send elapsed time is no more
than one-third of the prior-process aggregate for the same or a comparable
1,000-quote set using the same predeclared timestamp definitions and
aggregation method.
owner: sales
capability_boundaries:
- >-
EV_MULTI_007 is call-learning evidence: Gabriel Paunescu reports buyer
interviews rather than buyers answering directly in this authoring
task.
- >-
The reported context covers more than 25 production customers
collectively spanning M1_V1 through M1_V7 and one 1,000-quote batch.
Account identities, exact account descriptions, interview dates, and
calendar observation period were not supplied.
- >-
The represented language list was not supplied. This proof establishes
only the languages present in the sanitized evaluation package and does
not establish universal language coverage.
- >-
Upload, multilingual normalization, extraction and matching, and batch
intake are configured capabilities reported through buyer interviews.
MCP server d8189168 did not independently verify their implementation.
- >-
The 5 percent value is a buyer-accepted maximum per complete quote, not
an independently verified observed AI error rate. The required-field
list, truth-set construction, and mismatch-adjudication procedure must
be fixed before execution.
- >-
Exactly one independently reviewable result per source is required.
Missing results fail the batch acceptance criterion, but retry,
duplicate, timeout, and other partial-failure behavior are not otherwise
characterized.
- >-
MCP server d8189168 grounds the structured Create Quote surface and the
create-quote-document validation, formatting, and draft-storage path.
It does not establish a customer result.
- >-
The complete buyer-visible price-and-timing calculation was not verified
through MCP. Correct commercial values remain dependent on maintained
product, pricing, capacity, scheduling, and other master data.
- >-
"3x faster" is interpreted for this proof as aggregate start-to-send
elapsed time no greater than one-third of the prior-process aggregate
for the same or comparable 1,000-quote set. EV_MULTI_007 does not supply
the exact prior or observed durations, timestamp fields, comparison-set
construction, or aggregation method.
- >-
No raw database values, quote requests, timestamps, transcript,
recording, screenshot, export, or customer system was inspected during
authoring.
- >-
This proof does not measure universal language coverage, per-field
accuracy, correctness without review, capacity beyond one 1,000-request
batch, labor savings, implementation effort, retry behavior, duplicate
handling, price correctness, timing correctness, delivery performance,
win rate, margin, or physical-order shipping speed.
transfer:
level: prohibited
from_vertical_ids: []
conditions:
- >-
Use is limited to M1_V1, M1_V2, M1_V3, M1_V4, M1_V5, M1_V6, and
M1_V7. Any future or otherwise unlisted vertical requires separately
validated evidence and an explicit applicable canonical transfer
rule.
evidence_ids:
- EV_MULTI_007

Outstanding

  1. The safe evaluation package was not inspected during authoring. Before execution, fix the represented languages, required-field list, truth-set construction, source identifiers, and timestamp fields.
  2. The timing aggregation remains a declared test input. Use the same predeclared method for the prior and configured sets; the evidence does not supply exact baseline durations or the historical aggregation method.
  3. status: approved prevents approved generation. Lifecycle review remains required after execution and evidence review.