Defect Triaging Agent
We pull recent monocart-html reports from S3 and turn them into a reliability report for the Playwright suite, on demand.
Playwright · monocart-html · AWS S3 · TypeScript · LLM Analysis · Test Reliability
Teams struggle to see what automation reports are telling them over time. Most teams open the latest execution report and fix whatever failed. That gives us a view of the last run, but it does not show whether a spec has been flaky for weeks, keeps failing after retries, or is getting slower. We need more reports to find those patterns and make good decisions.
We built this agent to turn that report history into a clear reliability picture. Instead of comparing runs by hand, we can ask for the last 30 reports and see the patterns across the suite in one place.
We choose a window of 1 to 50 monocart-html reports from S3. If a report is incomplete or unreadable, we leave it out and tell the team how many runs we could use.
We parse the usable reports before involving the model. For every spec, we keep its attempt history, final outcome, duration, timestamp, and sanitized error details. The raw HTML stays out of the model context.
We group that evidence by spec and by day. The report separates final failures from retry-hidden flakiness, groups recurring errors, tracks duration changes, and ranks the tests that need attention.
Sending dozens of full monocart-html reports to the model would waste context and make the analysis less reliable. We parse each artifact first and pass the model only the compact, aggregated spec-level evidence it needs: attempt outcomes, final outcome, duration, timestamp, and sanitized error data.
A test that passes after retry can look healthy in a final-result-only view even when it is unstable. We retain the initial attempt and both retries, so the report separates flaky specs from specs that still fail after all retry attempts.
Complex tests can take longer than the rest of the suite and should not automatically become defects. We compare each spec with the suite baseline and its own daily history, then prioritize material slowdowns while keeping stable long-running specs as context.
A flaky spec fails on its first attempt and then passes on retry. A regularly failing spec still fails after both retries.
We can see how failures and durations change day by day, along with each spec’s average execution time. A complex test can stay slow without being treated as a defect; it rises in the queue when it gets slower.
We group similar errors with the affected specs and one sanitized example, so we can investigate the shared cause instead of reading the same failure message repeatedly.