This is a rubric, not a trophy case. Logos change monthly. Criteria do not. Use it to compare anything that claims to be an agent, copilot, or operator.
We do not take payment for placement. If that ever changes, About will say so first.
Weights (defaults — change them)
| Criterion | Default weight | What “5” looks like |
|---|---|---|
| Bounded actions | 5 | Allowlists, drafts vs prod, stop conditions |
| Trace | 5 | You can see tools, paths, URLs |
| Eval hooks | 4 | Tests, checklists, retries with a budget |
| Data control | 5 | Files you own; export; retention you understand |
| Caps | 4 | Spend, send, delete require a human or a hard limit |
| Latency / reliability | 3 | Boring, documented failure |
| Cost honesty | 3 | Meter visible before you burn it |
| Lock-in | 3 | Specs and packs leave when you do |
| Claims vs GA | 4 | Docs match the demo |
A hobby chatbot can score low on caps and still be fine if you never give it tools. The score is for the job, not a universal ranking.
How to run a 20-minute trial
- Write a tiny spec with two acceptance tests.
- Put a pack of three files in.
- Run once. Demand a trace.
- Fail a test on purpose. See if it retries or gaslights.
- Export everything. Close the account in your head: would you be stuck?
If you cannot complete that trial, the product is a demo.
Red flags (automatic cap at 2)
- “It just knows” with no permission screen
- No way to disable send/spend
- Memory you cannot delete
- Demo uses a human operator off-camera (it happens)
- Terms that train on your private packs with no toggle
- Scorecards on other blogs that are clearly affiliate sludge — including ours if we ever slip
What we will not score here
Individual model IQ leaderboards. They are noisy and already over-covered. We score the system around the model: files, tools, fences. That is the Bye Prompt beat.
Creative tools: Create By Prompt scorecard.
Payments rails: Pay By Prompt providers.
Shopping agents: Buy By Prompt (expanding).
Worked comparison (shape, not a league table)
When we compare two IDEs with agents, we would rather say:
- A: excellent trace, weak spend story (N/A if it cannot spend)
- B: pretty UI, no export of the pack, retry loop hidden
…than crown a winner. Your weights may invert ours (a solo student vs a shop that files SOC2).
Copy this table
Product:
Job:
Bounded actions (1–5):
Trace (1–5):
Eval hooks (1–5):
Data control (1–5):
Caps (1–5):
Reliability (1–5):
Cost honesty (1–5):
Lock-in (1–5):
Claims vs GA (1–5):
Notes / date:Date the row. Re-score when the product ships a real trace, not when they ship a new adjective.