← All posts

Measurement guide / 5‑minute read

How to tell whether a week of review-led fixes actually worked

Any consultancy can show you a week where revenue went up 8 per cent after they changed three things. Proving the three things caused it is much harder — here is what a single week can honestly demonstrate, and what it can't.

The standard version of this story goes: a restaurant read its reviews, ran three pre-shift fixes for a week, and finished 8.6 per cent up on revenue with 41 per cent fewer remakes. Sometimes there is a PDF.

You should not believe that story, and — this is the more useful point — you should not believe it about your own restaurant either. Not because the fixes don't work. They often do, and the ones that come out of guest language work more often than the ones that come out of a meeting. But a single week of revenue cannot tell you whether they worked, and a number presented as proof when it isn't one will cost you: you will keep a fix that did nothing, drop one that was working, and lose the ability to tell the difference next time.

So here is the honest version. Reviews are an excellent source of what to fix. Proving a fix worked is a separate discipline, it is not difficult, and almost nobody does it.

Why one week of revenue proves nothing

Try this before you read any case study, including one about yourself. Put your last eight weekly revenue totals in a column and look at the gaps between them.

Most venues find swings of several per cent in both directions, week to week, with no intervention at all. Weather. A bank holiday sitting in one week and not the other. Payday landing on a Thursday. One eighteen-cover booking. A rival down the road closing for refurbishment. School holidays starting. Against that background, a week that finishes up 8 per cent after you changed three things is entirely consistent with the three things having done nothing whatsoever.

The same problem sits inside average check. On roughly 800 tickets, a couple of large tables with wine will move your average more than a register upsell script will, and you cannot tell the two apart from the total. Attribution needs either a lot more time or a much narrower measurement.

Time is the expensive option. Narrowing is free.

Measure the thing you changed, not the thing you hope moves

The trick is to stop measuring the outcome you care about and start measuring the mechanism you actually touched. If you added a final bag check for sides, the question is not "did revenue rise" — it is "are sides still missing".

Three kinds of measurement survive contact with a single week.

Does the phrase still appear? You changed something because guests kept using particular words. Count how many times those words appeared in the two weeks before, then in the two weeks after. "Missing" and "forgot" going from eleven mentions to two is a real finding about a specific failure, and unlike revenue it is hard to explain away with the weather.

Counts of internal failures you already log. Remakes, comps, voids, refund requests from delivery partners, escalations that reached a manager. These are already recorded somewhere in your POS, they respond within days, and they are measured in whole events rather than percentages of a large number. They are the best short-horizon evidence you have.

Did the fix actually happen? The most common reason a fix shows no effect is that it stopped being done by Wednesday. Before concluding anything about the idea, check whether the checklist was filled in on Saturday night, which is the only shift where it was ever going to be hard.

Note what all three have in common: they are counts of events, they are within your control, and none of them require a model to interpret.

Write the number down before you change anything

The failure that ruins most of these attempts is not statistical. It is that nobody recorded the "before".

A baseline takes about twenty minutes and it has to be captured first, because after the fix your memory of the previous fortnight is worthless. For each thing you are about to change, write down: the exact guest phrasing you are targeting and how many times it appeared in a defined window; the internal count that should move if you are right; and what number would make you drop the fix rather than keep it.

That last one matters more than it looks. Deciding the kill threshold in advance is what stops a week of work from turning into a search for evidence that the week of work was worthwhile.

The arithmetic, using your numbers instead of mine

Once you have counts rather than percentages, you can put money against a fix honestly. Do it with your own figures — the point of the working below is the shape, not the values.

Say you log 80 remakes in a normal week. Take your own food cost for the dishes involved, add the labour to make them again, and add whatever proportion of those tables end up comped. Call the all-in cost per remake £C. If remakes fall to 55 for three consecutive weeks, the claim you can defend is 25 × £C a week — a number built from things you counted, not from a revenue line you hope to attribute.

Then be equally rigorous about the other direction. Vented lids cost something per unit. The pre-shift huddle costs ten minutes of several people's time, five days a week, and that is real money. A fix that saves less than it costs is worth knowing about, and you will only ever find out if you write both sides down.

This is a much smaller claim than "8.6 per cent revenue uplift". It is also one you can put in front of a sceptical business partner without flinching, and it compounds, because next quarter you can do it again and the baselines are already there.

Three ways a fix can look like it worked when it didn't

You started from a bad week. If you launched the fix the Monday after the worst weekend of the quarter, some of the improvement is simply the weekend not repeating. Compare against a normal fortnight, not against your reason for acting.

The complaint moved rather than stopped. "Missing sides" disappears and "wrong order" appears, because the bag check slowed the pass and something else gave. Always read the whole set of complaint categories after a change, not just the one you were hunting. A fix that relocates a problem reads as a success on a single-metric report.

Someone was watching. If the improvement coincides with the GM standing at expo all week, you have measured the GM, not the checklist. The test is week three, when everyone is back to their normal jobs. Fixes that only survive supervision are not fixes, they are surveillance, and they will quietly fail the moment attention moves elsewhere.

What to do with the answer

Three outcomes, three responses. If the count moved and stayed moved without anyone hovering, write the change into the standard and stop measuring it. If nothing moved, check whether the fix was actually performed before deciding the idea was wrong — those are different failures with different lessons. If you can't tell, the fix was scoped too broadly: pick a narrower one where the mechanism is a single countable event.

And if a fix cannot be verified within a fortnight by counting something you already record, that is a signal in itself. It may still be the right thing to do, but you are choosing it on judgement, not evidence, and it is worth being clear with yourself about which one you are doing.

Where this gets easier

None of this is hard. It is a spreadsheet and a habit. The reason it rarely happens is that the "count the phrase" step lands on the same person who is short-staffed on Saturday, and it is the first thing to go.

That is the part OMMU takes on. It reads every review as it lands, sorts the complaints into the eight areas of your business so the same phrase is counted the same way every week, and puts what changed into one short email — including the line that matters most here: the complaint that has stopped appearing, which is your evidence that a fix took hold. The sample report is a real briefing sent to a multi-location owner, if you would rather see the format than read about it.

Whether or not you use us, capture the baseline before you change anything this week. It is twenty minutes, and it is the difference between knowing your fixes work and having a story about them.