Onlytool
Automated conversations and selling workflows.
Visit OnlytoolChatbot Lab / Evaluation
Great demos are easy. Reliable behaviour takes evidence. Score the things that matter and make your next trial count.
Weighted score
/ 100Illustrative inputs. Change a score to see the method.
Full scorecard: six dimensions, editable weights and an exportable record.
A USEFUL NEXT STEP
THE DECISION, EXPLAINED
Use a fixed scenario set, an explicit scoring rubric and separate mandatory pass conditions. A fluent answer can still use the wrong context, make an unsupported promise or continue after an operator pauses the workflow. The scorecard helps organise your observations; it does not supply them.
Evaluate the same configuration on the same examples before comparing totals. Keep untested capabilities marked as unknown. A weighted score is useful for trade-offs between acceptable candidates, but a failed mandatory control should remain visible regardless of the average.
See it in a worked exampleUse a fixed set of scenarios.
Weight what matters to you.
Keep hard failures visible.
| Dimension | Observe this | Keep this evidence |
|---|---|---|
| Context | Uses the correct authorised facts | Scenario, available context and output |
| Voice consistency | Follows the brief without inventing facts | Brief version and departures |
| Controls & handover | Pauses and routes exceptions as specified | Event order and assigned owner |
| Billing & data | Explains the actual fee and access lifecycle | Written quote and data answers |
FROM READING TO DOING
Repeatable scenarios for context, corrections and stop behaviour, with an exportable test record.
BUILD YOUR SHORTLIST
Explore documented options for your workflow.
Onlytool is our own product.
Automated conversations and selling workflows.
Visit OnlytoolAI chat and agency CRM.
Visit SubstyAI chatting and creator management tools.
Visit SupercreatorPublic provider descriptions checked 2026-09-10. Listings are not independent performance ratings. Read the method.
GO ONE LEVEL DEEPER
Run the same scenarios against every candidate and keep a record of the evidence.
Read the complete guideGOOD QUESTIONS
Clear answers, useful context.
No. It is your assessment of your own observations, with weights you control. We do not assign these scores to providers.
Only that the entered scores produced that weighted result. It does not establish safety, policy compliance, reliability or revenue performance.
Use the download button to save a text record of the weights and scores. Keep supporting evidence in your own approved system.