What is this product
The "Testbed" is developed by Invent LLC, a company within the RESTART Group. The testbed is deployed within the customer's environment and evaluates the AI system along two axes simultaneously: whether it can handle the specified load and whether responses remain correct as the load increases. Both axes are measured using the same traffic within a single test run.
The main output is not a web page with charts, but a test report in PDF and DOCX formats: containing the test program and methodology, test conditions, results for each axis, pass/fail verdicts for each acceptance criterion, and a signature block. The document is intended to be included in the delivery acceptance certificate package.
What one can feel
RAG assistants
Corporate knowledge bases and document assistants are evaluated both on response time and whether the responses are grounded in sources as the number of queries increases.
Support chatbots
Customer-facing and internal bots ahead of the seasonal peak: how many concurrent requests can they handle, and do they start making up facts?
LLM API and inference services
Own inference pipeline: throughput, queue formation time, behavior upon node failure.
Assistants in business processes
Scenarios where the model's response continues into the process—there, an error is costlier and is verified separately.
How are acceptance criteria defined
The test profile is agreed upon prior to execution and becomes the single source of truth: what is recorded in it is quoted verbatim in the protocol. Typically, four groups of parameters are captured.
| What are we setting | Parameter examples | Who is involved in the approval |
|---|---|---|
| Load profile | Single requests, one-time burst, stepwise ramp-up, background load with pauses, random arrivals within a time window | AI Product Owner, Site Reliability Engineer |
| Validation dataset | A set of real questions and reference answers, sample size for quality assessment | Process owner, subject matter experts |
| Acceptance thresholds | Success rate, P95 response time, allowable number of failures, minimum quality score, number of users without queue | Acceptance committee, test customer |
| Run boundaries | Environment, system version, execution window, access to infrastructure metrics | IT Operations, Security Service |
The infrastructure metrics section is optional: even without access to containers and GPUs, the test stand still loads the system via HTTP and delivers a verdict based on the remaining criteria.
What does the customer get?
- Test protocol in PDF and DOCX formats based on an editable template: title page, program and methodology, conditions, load and quality results, verdict for each criterion, and a signature block.
- Response for each acceptance criterion in the format "threshold — actual — PASS or FAIL", without general statements such as "there are comments".
- Example applications: question, system response, evaluation, and textual explanation of the verdict — the evaluation logic can be manually rechecked.
- Behavior under load data: queue formation stage, concurrency level at which rejections begin, queue and inference engine memory state, if metrics access is provided.
How it differs from audit and laboratory
The group's three proposals are easily confused—they address different questions and are applied at different stages.
| Format | What question does it answer? | What's the output |
|---|---|---|
| Secure AI audit | What risks does the AI loop architecture entail, and how are data and access managed within it? | Risk register, architecture review, remediation roadmap |
| Information Security Laboratory | How a specific security control or product behaves within the customer's architecture | Pilot protocol, compatibility matrix, implementation requirements |
| Polygon | Does the AI system pass acceptance testing according to agreed-upon load and quality criteria? | Test protocol with PASS/FAIL verdict for the acceptance certificate |
Bench tests are not certification and do not confirm compliance with regulatory requirements—they confirm fulfillment of the criteria agreed upon by the acceptance parties.
How to get started
Application and call
We discuss the system type, available traffic of real queries, and access to infrastructure metrics.
Test Profile
We'll align the load profile, reference dataset, and acceptance thresholds with your system.
Run in your environment
The test bench stresses the target and simultaneously evaluates the response sample using a local judge.
Protocol
You receive a document with a verdict for each criterion, ready to be included in the set with the report.
Request a demo or pilot
We will show the product on your environment and offer a pilot format with a clear output result.
