Resources

Everything the bird has brought back from the mine.

The public layer of Coalbird is here to make the category legible: what agents can do, where stores break, and how to price the work of watching that channel.

Benchmark

The Hundred

A public league table of 100 stores the bird has walked into, with pass/fail outcomes for product findability, returns policy, and add-to-cart.

Browse the benchmark
Research

Agentic Commerce Readiness Report

69% of 100 well-known DTC stores failed when an AI shopping agent tried to add a product to the cart (Coalbird Agent Audit, 2026-08).

Read the report
Field notes

The bird tweets

Short traces, quotes, and failure patterns from real agent runs.

Read field notes
Working notes

How to make sense of agent failures.

How to think about agent-experience testing

Start with missions a shopper would delegate: find, compare, choose, add, verify, and understand policy. Then grade only outcomes the agent can prove.

What to inspect before the bird goes in

Variant selectors, popups, region gates, cart state, search relevance, policy pages, product data, and anything important hidden inside images.

How to use a Coalbird trace

Treat it like session replay for agents: the trace shows the step, the screenshot shows what the agent saw, and the verdict tells you why the journey stopped.

When monitoring matters

Weekly runs are enough for stable stores. More frequent reruns make sense around launches, seasonal promos, template changes, and international expansion.

Trial

Ready to test your own store?

Start with the native intake. The final step is a Calendly call where we scope the trial missions and decide what the bird is allowed to touch.

Start trial