Methodology

How we test

The full method, published so you can check it — six stages, fixed test sets per category, and a score with its working shown.

Method version 1.0 · Last updated August 2026

The principle

A review is only worth something if the next reviewer could run the same test and get the same answer. That means the test has to be fixed before the product arrives, not designed around what the product happens to be good at. TanLab's test sets are written per category and applied identically to every product in that category, cheap or flagship.

Three rules hold the whole method together:

  • Same conditions. Comparisons are shot in the same session, same light, same temperature, same network. A phone tested in December against one tested in June is not a comparison.
  • Three runs minimum. Every measured number is the median of at least three runs. One run is an anecdote.
  • Publish the raw numbers. The measurements go in the video and in the written review, not just the conclusion drawn from them.

The six stages

01 · Unbox

The unboxing is recorded untouched — first power-on, factory firmware, factory settings. We log the serial and firmware version, because a firmware update mid-test invalidates earlier numbers and has to be declared. Build quality gets a close inspection here: panel gaps, blade grind, thread quality, coating, the state of the accessories. We also check what is genuinely in the box against what the retail listing promised, which is a surprisingly frequent miss on budget gear.

02 · Real usage

The product becomes the daily driver for a minimum of fourteen days. Phone goes in the pocket as the only phone. Knife does the actual cutting. Flashlight is the one on the keys. Throughout, a dated log records every irritation, failure and pleasant surprise — this is where the "spec sheet says great, life says otherwise" problems show up, and no bench test finds them.

03 · Measured performance

The fixed test set for the category runs under controlled conditions. Settings are recorded and reused for every product in that category. Ambient temperature is noted for anything thermal. Nothing here is a synthetic benchmark score quoted alone — a benchmark number without the sustained-load result next to it tells you almost nothing.

04 · Stress test

The product gets pushed past comfortable: sustained load until it throttles, runtime until it shuts down, water exposure to its claimed rating (never past it, unless the claim is the thing being tested), and drops from realistic carry height onto a realistic surface. The point is not spectacle. The point is to find the failure mode before you do, at your expense.

05 · Head-to-head

The closest rivals at the same price run the identical test set, back to back, in the same session. A verdict that isn't relative to alternatives isn't useful — "good battery life" only means something against what else that money buys today.

06 · Verdict

Scores are weighted into the final number, the working is published, and the review closes with the two sentences that matter most: who should buy this, and who should not. A product can score 91 and still be the wrong purchase for most people, and the review has to say so.

The TanLab Score

One number out of 100, built from weighted category scores. The weights differ per product type, because battery life matters more on a phone than on a pocket knife, and they are published so you can disagree with them intelligently.

Smartphones

CategoryWeightWhat decides it
Camera30%Identical-scene comparison, day and night, stills and video, across all lenses
Battery25%Real-day screen-on time plus a controlled loop test, and the charging curve
Performance20%Sustained load and thermals, not peak benchmark; app-open and gaming stability
Build & display15%Measured brightness, panel quality, materials, ingress rating, drop result
Value10%Street price against what the rivals deliver on the same day

Flashlights

CategoryWeightWhat decides it
Output honesty30%Measured output at 30 s and 10 min against the claim on the box
Runtime25%Time to 10% of initial output, per mode, at room temperature
Beam & throw20%Distance, spill quality, tint and artefacts on a measured wall shot
Build & durability15%Water rating verified, drop test, threads, switch feel, charging port
Interface & value10%Can you get the mode you want in the dark, half-asleep — and for the price

Knives & EDC

CategoryWeightWhat decides it
Edge retention30%Sharpness measured before and after a fixed cut count on identical media
Build & lock25%Grind consistency, centring, blade play, lock strength and spine-whack behaviour
Carry20%Weight, clip design, deployment, pocket comfort over two weeks of real carry
Steel & finish15%Corrosion after controlled exposure, ease of resharpening, coating wear
Value10%What the same money buys elsewhere in the same steel and build

Drones, cameras, wearables and audio

Each has its own published weighting on the same principle — the thing you actually bought the product for carries the most weight. Drone scoring leads on real flight time and signal stability; camera scoring leads on image quality and autofocus reliability; wearable scoring leads on sensor accuracy against a reference; audio scoring leads on tonal balance and ANC depth. The per-category tables are published with the first review in each category.

What the numbers mean

ScoreVerdictMeaning
90–100Lab ApprovedClass-leading. Buy it without researching further.
80–89RecommendedStrong, with known trade-offs that are spelled out in the review.
70–79SituationalRight for a specific buyer, wrong for most. Read who.
60–69CompromisedSomething important is weak. Only at a discount.
Below 60AvoidThe claim and the reality are too far apart.

What voids a review

A test result is thrown out and re-run if any of these happen:

  • Firmware changed mid-test in a way that affects the measured behaviour
  • Ambient conditions drifted outside the recorded range during a thermal or runtime test
  • The unit shows a defect that isn't representative — a faulty sample gets replaced, not reviewed
  • A comparison couldn't be run in the same session against the same conditions

A review is not published at all if the product cannot be tested independently — for example if a brand will only supply a unit on the condition of copy approval. That condition is always declined.

Corrections

If a measurement turns out to be wrong, the correction is added to the top of the written review with the date and what changed, and pinned in the video comments. Scores can move after a long-term follow-up; when they do, the original number stays visible next to the new one. Spot an error? Tell us — corrections from viewers are welcome and get credited.

See it in practice

The method is only as good as the footage that backs it. Every test in this document is shown running on the TanLab YouTube channel.