Guide · evidence quality
False positives in AI website testing: telling a bug from a model guess
An AI tester that looks at screenshots can be confidently wrong: it can call readable text low contrast, miss a label that is there, or flag a design preference as a defect. A finding is only useful if you can check it. This guide shows what to look for in a finding, and how Sybqa separates measured facts from model opinion.
Facts checked
What makes a finding checkable
Whatever tool produced it, a finding you can act on has four parts: where it happened (route, viewport and state), what was observed, the steps to reproduce it, and a screenshot. If a finding lacks any of these, treat it as a suggestion, not a defect.
| Kind | Example | How to treat it |
|---|---|---|
| Measured fact | Text contrast is 3.2:1 where WCAG asks for 4.5:1 | Verify by measuring. If the number holds, it is a defect. |
| Deterministic rule | A visible undefined, or a 20 px tap target | Reproducible by anyone. Check the rule matches your intent. |
| Observed journey result | The form submitted and no confirmation appeared within five seconds | Reproduce with the listed steps; the outcome is observable. |
| Model claim, confirmed by a measurement | A model says a label is missing and the page confirms no accessible label | Treat as measured. |
| Model claim, not measurable | The call to action feels hidden | A lead to review with your own eyes, not a failure. |
How Sybqa handles model claims
Sybqa keeps browser observations, model suggestions, source inspection and real provider outcomes as separate kinds of evidence. With AI review on, runs that include journeys have each distinct screen reviewed by two independent vision models by default, and their claims are checked against measurements taken in the page.
- A claim about contrast, labels or target size that the measurement confirms becomes a failure with the measured value.
- A claim that the measurement refutes is dropped, and the report summary lists what was refuted or dropped.
- Claims that cannot be measured are scored and, if they pass the threshold, shown as review leads, not failures. Uncertain high-severity claims get one further opinion.
- Deterministic checklist rules, such as touch-target size and persistent labels, do not depend on a model at all.
What this does not establish
Review any finding in a minute
- Open the screenshot and find the thing the finding names. If you cannot find it, say so; the finding is wrong or unclear.
- Repeat the steps to reproduce at the stated viewport.
- If a number is quoted, check it, for example with your browser’s accessibility inspector.
- Decide: fix, ignore with a reason, or mark for a person with the right context.
Common questions
Will Sybqa report a problem only a model noticed?
It can, but as a lead to review rather than a failure. Failures come from measurements, deterministic rules and observed journey results.
Can I see what was dropped?
Yes. The review summary in the report lists claims that were refuted by a measurement or dropped.
Does AI review always run?
No. The free plan runs deterministic checks without AI review; AI review is part of the paid plan and uses credits.
Sources
Related
Try Sybqa on your site
Paste a link, review the plan, and read the evidence report. The free plan gives 10 runs a month without AI review after you sign up.