Review Methodology

How we score software

Our Review Methodology

We evaluated each platform using a structured, criteria-based scoring system rather than general impressions — 95 individual items across 5 categories, covering the core workflows a field service business runs through every day.

What we tested

We created live accounts with each platform and worked through the same set of core workflows in every tool — creating jobs and clients, viewing job details, tracking time, collecting payments and signatures, managing photos and attachments, navigating to job sites, and more. Every screen was captured directly from our own accounts, and every item was tested hands-on rather than assumed.

How we scored

Each item was scored on a 1–10 scale (in 0.1 increments) using three consistent questions:

  • Create — How complicated is it to set this up for the first time?
  • Edit — How easy is it to go back and change what you just made?
  • Detail — What useful information does the tool capture, and what’s missing or annoying?

A mid-range score means the feature works as expected with no particular friction or standout quality. The top of the scale is reserved for best-in-class execution — fast, obvious, and doing more than we expected. The bottom reflects a feature that’s broken, painful, or effectively missing. Where a feature genuinely didn’t exist in a platform, we marked it N/A and excluded it from averages rather than penalizing a tool for a feature it doesn’t offer.

Two independent passes, not one

We ran every item through two separate scoring passes and kept them distinct rather than blending them into a single number:

  1. Hands-on pass — Our own team, testing live and creating real data in each platform, screen by screen.
  2. AI-assisted pass — A separate score generated using only information available from each vendor’s own site and public discussion forums, with no access to our hands-on results.

Wherever the two passes landed more than 0.2 points apart on an item, we flagged it and did a manual audit to determine which score was accurate and why the gap existed, rather than simply averaging the difference away.

This two-pass structure is also why we don’t lean on aggregated review-site scores the way a lot of comparison sites do. We’re not re-packaging other people’s star ratings — every score reflects work we did ourselves, cross-checked against a second independent method, with a documented reason any time the two disagreed.

Keeping it consistent

To avoid favoring one platform over another, we applied the same lens and the same audit process to every tool in the comparison, and periodically checked whether any recurring score gaps reflected real, observed differences rather than inconsistent standards on our part.