Review operations

Plan review usage around a question

Treat AI reviews as a counted, bounded resource and spend each one on a question worth settling.

The new review form is open: a website selected, a URL pasted, a goal typed into the box labelled "What do you want to learn?", and a line of fine print under the submit button reading "Uses 1 review · 61 remaining". The person filling it in has already started two runs this week, both against the pricing page, both with goals that differ by a handful of words, and both reports are still sitting unopened in the reviews list.

The instinct behind that is a reasonable one: a review looks like a button, and a report that did not quite land seems like something you fix by pressing again with better wording. The counter disagrees, and it disagrees precisely. The allowance check is a strict less-than, so the last review you may start is the one that takes the count up to the limit and the next request is refused outright; and the count is incremented in the same database transaction that creates the review, with the usage row locked while it happens. The charge lands when you press the button, not when a report arrives.

What you are spending is also smaller than it looks from outside. A run is bounded before it begins: the reviewer reads at most four pages, makes at most five model calls, refuses any call that would carry its generated output past six thousand tokens, and gives up entirely at a two hundred and thirty second deadline. Pressing the button a second time does not make that envelope any larger.

The allowance is therefore not a throttle bolted onto an open-ended machine. It is a count of small, bounded readings of public page text, and the unit worth planning around is not the run but the question the run is meant to settle.

Two colleagues at a desk, one writing a question, one at a laptop
Spend a review on a decision you are ready for.

Count The Allowance Before You Frame The Question

There are only four sets of numbers to know. Focus carries thirty AI reviews and ten thousand visitor submissions, Growth carries one hundred reviews and fifty thousand submissions, Portfolio carries three hundred reviews and two hundred thousand submissions, and the fourteen day trial carries three reviews and five hundred submissions. Thirty reviews on Focus works out at roughly one per working day if you spread them evenly across a month, which is a useful way to feel the size of the allowance before half of it has gone in a single afternoon.

The common assumption is that a published limit is a soft target the product will quietly stretch and bill for afterwards. It does not stretch. When the step that starts a review finds the locked meter already at the plan's review count, it refuses the request and answers "You've used your AI review allowance. Your existing reports remain available." The billing panel in workspace settings states the same policy in its own words: "Monthly allowance. No surprise overages." Reaching the limit stops new reviews and changes nothing about the reports you already hold.

Because the count goes up when the request is accepted rather than when a report finishes, the number in front of you is a count of starts, not of finished work. The overview prints it as a plain subtraction of usage from the plan allowance, described as the reviews left in your allowance, and the review form repeats the same arithmetic beside the words "Uses 1 review". Read that number before you write the goal rather than after you have already decided you want another run.

Treat The Trial Meter As A Total, Not A Monthly Rate

Usage is recorded against a period, and the period is decided by one rule with two outcomes. A paid workspace uses the current year and month in UTC, so its meter is a row named for this month; every other workspace uses a single fixed period named EVALUATION. On the first of the next month a paid workspace begins writing into a row that did not exist before and is created empty, which is the whole of the reset: no scheduled job clears anything, the name of the row simply changes. That is why the billing panel can promise "Resets on the first of the month, UTC" with no machinery standing behind the promise.

A workspace that is not paying never changes rows. Every review and every visitor submission it records lands in that same EVALUATION period for the entire life of the trial, so three reviews means three in total rather than three per month, and five hundred submissions means five hundred in total. The billing panel says exactly this where the monthly wording would otherwise sit, printing "500 total during your trial" beneath the submissions bar and the trial end date beneath the reviews bar. The pricing page puts both halves in one clause: "Paid allowances reset on the first of each month in UTC; the trial allowances are totals."

So a trial is not a small month and should not be planned as one. You have three runs to spend across fourteen days, after which a new review is refused with "Your evaluation has ended. Your feedback stays here. Choose a plan to run more reviews." Three runs is enough to read one page through three different lenses, or to read one page before a change and again after it has shipped. It is not enough to absorb a rephrased repeat, and no first of the month is coming to hand the decision back to you.

A five-step flow from a written question through the usage meter into the review allowance, contrasting the paid monthly period key that resets with the single trial period key that does not.
A trial allowance is a total, not a monthly count.

Spend The Run On A Question The Budget Can Answer

The form gives you two levers and no others: a goal of between ten and fifteen hundred characters, prefilled with "Help a first-time visitor understand the product and confidently take the next step", and one of four lenses named First-time buyer, Product leader, Usability reviewer and Inclusive content reviewer. Both levers travel into every model call the run makes, including the call that chooses where to go next, so the goal is not a label printed on the report; it is the instruction that decides which four pages get read at all.

The instinctive goal is a broad one, on the theory that a wider question returns more. It returns less, because the width is spent on navigation instead of on reading. After each page the reviewer offers the model at most thirty of the links it observed and asks for one of them by position, or for a decision to stop, and the loop ends at four pages either way; each page then contributes only its first six thousand characters, cut into excerpts of at most four hundred and fifty characters, and a finding that does not name one of those excerpts is discarded by the server before you ever see it.

Knowing the rest of the envelope explains why a second attempt at the same broad question comes back just as thin. Five model calls is the ceiling for the whole run, and a complete run typically spends three of them choosing the next page, one writing the report, and one on a second pass that acts as an independent skeptical evidence editor; any call whose own limit would carry the run's generated output past six thousand tokens is refused rather than trimmed. The report request asks for up to five directly supported improvements and states plainly that zero findings is a valid answer. That is the whole of what one unit of your allowance buys.

Diagram of one question, a bounded review and selected follow-up

Know Which Failures Return The Credit

Not every stopped run costs you a review. A run that did not complete and never reached its first model call gives the review back to the meter, and it gives it back to the period the review was charged to rather than to whichever period happens to be current when the failure is recorded. The first model call is the boundary, and it is a sharp one: a run that fails while fetching the starting page, or that you cancel while it is still waiting in the queue, returns the review, while a run that reaches the model and then fails does not.

Three guards sit in front of the meter and are easy to mistake for the allowance itself. A workspace may have only two reviews in flight at once, and a third attempt is answered with "Two reviews are already running. Wait for one to finish before starting another." Review starts are rate limited per workspace at twelve per hour and sixty per day, whatever the plan. And every start carries an identifier created with the form, so a resubmission returns the review already saved under that identifier: an impatient second click on "Start the review" costs one review, not two.

Runs that get stuck are swept rather than left to occupy one of the two slots. A review that has waited in the queue for more than six hundred seconds, or that has run past its lease, is closed as failed with "This review timed out. You can start a new review." and the same credit rule applies unchanged. That is the reason to look at a slow review's actual state before restarting it: a run that has already reached the model has already been counted, and starting a replacement adds a second charge rather than cancelling the first.

What This Does Not Tell You

A spent review buys an interpretation of public page text and nothing more. The limitations stored on every completed report say so exactly: "Public-page content review. Visual layout, live interactions, page speed, accessibility compliance, and logged-in workflows were not tested." The report page repeats the point above the findings: "These are AI interpretations, not accounts from real visitors. Use them to form better questions and decide what to test." No allocation of the allowance changes that; thirty runs on Focus are thirty readings of text, and not one of them observed a visitor, submitted a form or looked at a rendered page.

The meter is equally silent on whether a run was worth starting. A review that completes with nothing surviving the evidence check is counted in full, and it stores the summary "This limited review found no sufficiently supported change to recommend. That does not mean the site is flawless. Try a more detailed public page or a different review goal." The usage the product reports for a workspace is two integers, reviews and submissions for the current period. It does not record which findings you accepted, which you rejected, or whether anyone opened the report; tracking a finding does create an AI-sourced item in the inbox, so that decision leaves a trace, but nothing joins the trace back to the allowance.

First Steps

Three actions turn the allowance from a number you notice at the limit into a constraint you plan around, and each of them uses a surface that already exists in the product.

  1. Open workspace settings and read both usage bars together with the line printed underneath them: a paid workspace shows "Resets on the first of the month, UTC", while a trial shows its end date and "500 total during your trial". Write down the remaining review count and, on a trial, the date the fourteen days expire.
  2. Write the goal as one question about one audience and one decision, inside the ten to fifteen hundred character field, and pick the lens that matches that audience instead of accepting the prefilled first-time buyer goal. Start from the URL the question actually concerns, because the run reads at most four pages outward from there.
  3. Before starting another run, open the last report, track the findings you accept into the inbox, and write down which ones you reject and why. Start the next review when a change has shipped or a genuinely different question has appeared, not when a report disappointed you.

Let The Next Question Earn Its Review

The two halves of this fit together more tightly than they first appear. The per-run budget is deliberately small, at four pages, five model calls, six thousand tokens of generated output and two hundred and thirty seconds, because a narrow reading carrying verified quotations is worth more than a broad one without them; and the per-period allowance is a count of exactly those small readings. A workspace that presses the button whenever a report disappoints turns a bounded, evidence-checked instrument into a slot machine, and the meter records the attempts either way.

So let the question set the pace. On a trial that means three questions and no rewrites, because the EVALUATION period never rolls over and nothing about the calendar will rescue a wasted run. On Focus, Growth or Portfolio it means thirty, one hundred or three hundred questions in a UTC month, refreshed because the name of the period changed rather than because anyone promised it would. The counter under the button is not an obstacle to the work; it is the closest thing the product keeps to a record of how many things you decided were worth finding out.