The principle
A review is only worth something if the next reviewer could run the same test and get the same answer. That means the test has to be fixed before the product arrives, not designed around what the product happens to be good at. TanLab's test sets are written per category and applied identically to every product in that category, cheap or flagship.
Three rules hold the whole method together:
- Same conditions. Comparisons are shot in the same session, same light, same temperature, same network. A phone tested in December against one tested in June is not a comparison.
- Three runs minimum. Every measured number is the median of at least three runs. One run is an anecdote.
- Publish the raw numbers. The measurements go in the video and in the written review, not just the conclusion drawn from them.
The six stages
01 · Unbox
The unboxing is recorded untouched — first power-on, factory firmware, factory settings. We log the serial and firmware version, because a firmware update mid-test invalidates earlier numbers and has to be declared. Build quality gets a close inspection here: panel gaps, blade grind, thread quality, coating, the state of the accessories. We also check what is genuinely in the box against what the retail listing promised, which is a surprisingly frequent miss on budget gear.
02 · Real usage
The product becomes the daily driver for a minimum of fourteen days. Phone goes in the pocket as the only phone. Knife does the actual cutting. Flashlight is the one on the keys. Throughout, a dated log records every irritation, failure and pleasant surprise — this is where the "spec sheet says great, life says otherwise" problems show up, and no bench test finds them.
03 · Measured performance
The fixed test set for the category runs under controlled conditions. Settings are recorded and reused for every product in that category. Ambient temperature is noted for anything thermal. Nothing here is a synthetic benchmark score quoted alone — a benchmark number without the sustained-load result next to it tells you almost nothing.
04 · Stress test
The product gets pushed past comfortable: sustained load until it throttles, runtime until it shuts down, water exposure to its claimed rating (never past it, unless the claim is the thing being tested), and drops from realistic carry height onto a realistic surface. The point is not spectacle. The point is to find the failure mode before you do, at your expense.
05 · Head-to-head
The closest rivals at the same price run the identical test set, back to back, in the same session. A verdict that isn't relative to alternatives isn't useful — "good battery life" only means something against what else that money buys today.
06 · Verdict
Scores are weighted into the final number, the working is published, and the review closes with the two sentences that matter most: who should buy this, and who should not. A product can score 91 and still be the wrong purchase for most people, and the review has to say so.
The TanLab Score
One number out of 100, built from weighted category scores. The weights differ per product type, because battery life matters more on a phone than on a pocket knife, and they are published so you can disagree with them intelligently.
Smartphones
| Category | Weight | What decides it |
|---|---|---|
| Camera | 30% | Identical-scene comparison, day and night, stills and video, across all lenses |
| Battery | 25% | Real-day screen-on time plus a controlled loop test, and the charging curve |
| Performance | 20% | Sustained load and thermals, not peak benchmark; app-open and gaming stability |
| Build & display | 15% | Measured brightness, panel quality, materials, ingress rating, drop result |
| Value | 10% | Street price against what the rivals deliver on the same day |
Flashlights
| Category | Weight | What decides it |
|---|---|---|
| Output honesty | 30% | Measured output at 30 s and 10 min against the claim on the box |
| Runtime | 25% | Time to 10% of initial output, per mode, at room temperature |
| Beam & throw | 20% | Distance, spill quality, tint and artefacts on a measured wall shot |
| Build & durability | 15% | Water rating verified, drop test, threads, switch feel, charging port |
| Interface & value | 10% | Can you get the mode you want in the dark, half-asleep — and for the price |
Knives & EDC
| Category | Weight | What decides it |
|---|---|---|
| Edge retention | 30% | Sharpness measured before and after a fixed cut count on identical media |
| Build & lock | 25% | Grind consistency, centring, blade play, lock strength and spine-whack behaviour |
| Carry | 20% | Weight, clip design, deployment, pocket comfort over two weeks of real carry |
| Steel & finish | 15% | Corrosion after controlled exposure, ease of resharpening, coating wear |
| Value | 10% | What the same money buys elsewhere in the same steel and build |
Drones, cameras, wearables and audio
Each has its own published weighting on the same principle — the thing you actually bought the product for carries the most weight. Drone scoring leads on real flight time and signal stability; camera scoring leads on image quality and autofocus reliability; wearable scoring leads on sensor accuracy against a reference; audio scoring leads on tonal balance and ANC depth. The per-category tables are published with the first review in each category.
What the numbers mean
| Score | Verdict | Meaning |
|---|---|---|
| 90–100 | Lab Approved | Class-leading. Buy it without researching further. |
| 80–89 | Recommended | Strong, with known trade-offs that are spelled out in the review. |
| 70–79 | Situational | Right for a specific buyer, wrong for most. Read who. |
| 60–69 | Compromised | Something important is weak. Only at a discount. |
| Below 60 | Avoid | The claim and the reality are too far apart. |
What voids a review
A test result is thrown out and re-run if any of these happen:
- Firmware changed mid-test in a way that affects the measured behaviour
- Ambient conditions drifted outside the recorded range during a thermal or runtime test
- The unit shows a defect that isn't representative — a faulty sample gets replaced, not reviewed
- A comparison couldn't be run in the same session against the same conditions
A review is not published at all if the product cannot be tested independently — for example if a brand will only supply a unit on the condition of copy approval. That condition is always declined.
Corrections
If a measurement turns out to be wrong, the correction is added to the top of the written review with the date and what changed, and pinned in the video comments. Scores can move after a long-term follow-up; when they do, the original number stays visible next to the new one. Spot an error? Tell us — corrections from viewers are welcome and get credited.
See it in practice
The method is only as good as the footage that backs it. Every test in this document is shown running on the TanLab YouTube channel.