Evaluating health data providers

The cheaper quote moves the cost to your roadmap.

Health data platforms sell two things. First the raw feed from phones and wearables. Then the analytics layer above it: scores, biomarkers, archetypes, tags and insights. Most price separately, so a cheaper quote is often not a quote for the same product.

Sahha is a health data API that includes both. This page sets out what the second layer costs to build, using our own numbers, so you can compare quotes at equal scope.

to build the analytics layer
15 months
before modelling can start
8 months
for the analytics layer
$1.2M in salary
a first-generation pipeline
10x to run
Before comparing price

Check what the quote covers.

Is scoring included, or an add-on?

Other health APIs price raw access and derived analytics separately. A scoring add-on can roughly double an entry-level quote, so a cheaper headline number may not cover what you are buying the platform for.

Sahha Included at every paid tier. Scores, biomarkers, archetypes, tags and insights ship together.

What happens when a user has no wearable?

Scores built on physiological inputs return null without them. A quote that assumes wearable coverage is a quote for the subset of your users who own one.

Sahha Every score works phone-only. A wearable adds signal, it is never required.

How long before a new user sees anything?

Baseline requirements of 7, 14, 28 or 30 days are common before a score computes at all. That window is dead time in your onboarding, every time.

Sahha Immediate in almost all cases, and calibrated from day one, because backfill supplies the history.

Who owns validation, and the liability?

Customizable scoring sounds like flexibility. In health it means the validation, and the liability for an unvalidated number someone changes their life on, transfers to you.

Sahha Validated against our own collected data, by researchers who publish it.

Provider-by-provider detail is in the comparisons. What is included at each tier is on pricing.

If you build it instead

Fifteen months, on the fastest path there is.

Critical path

15 months
Ethics protocol 2 mo
Data collection 6 mo
Modelling and scores 7 mo

8 months before a model can be trained. Headcount does not shorten it.

Running alongside

Four generations, 24 months
First build Batch, not streaming
Rebuild Streaming · 10x to run
Rebuild Retooled · 4x to run
Current Rearchitected · 1x to run
15 months of critical path ends here
Phase Duration Budget shortens it?
Ethics protocol A commercial review board, which is the fast option. Ours took closer to four through an academic one. 2 months No
Longitudinal data collection The aggressive floor, and it costs you a full seasonal cycle. We collected for nine to twelve. No number of hires shortens either. 6 months No
Pipeline, across 3 to 4 rebuilds Each rebuild took roughly six months. No first plan contains four of them. runs alongside Partly
Discarded model approaches The approaches that do not work, before the one that does. Ours took three to four months. 2 months Partly
Scores Two to three weeks each once the data and the approach exist. Readiness took a month. ~3.5 months Yes
Archetypes Built on the same foundation. 1.5 months Yes
Insights, trends and comparisons The easy part, and the part most teams imagine is the whole job. ~2 weeks Yes

Fifteen months only holds if you take all four of these. We took longer on each.

Commercial IRB instead of an academic board
Realistic, and genuinely faster. Costs money rather than time, and needs someone who can write the protocol, which means knowing what you intend to collect and why.
Six months of collection instead of nine to twelve
Captures half an annual cycle. Sleep, activity and mood all move with season and daylight, so a single-season dataset carries a bias you will not detect until you deploy into the other half of the year.
Fewer repeated clinical measures per participant
Thinner validation, most visibly for anything mental-health related, where the label is the expensive part.
Validate fewer scores
The fastest real saving. It also means shipping less than you set out to build.

What it costs in people

These are our own figures, from building this layer ourselves: ethics protocol, longitudinal collection, models, and four generations of pipeline. Scoped to what you would replicate, so front-end and dashboard work is excluded. Rates are published benchmarks, not our payroll, so the model stays checkable.

Role FTE Person-years Cost
ML researchers (PhD) Modelling cannot begin until collection completes. Costed as hired late, the cheaper of the two options. 2 1.5 $318,033
Pipeline engineers 2 4 $720,000
Senior infrastructure Part allocation. The rest goes to API, webhooks and dashboard work you would not replicate. 0.5 1 $180,000
Total 6.5 $1,218,033
If you build the raw feed too Getting the data off phones and out of wearable accounts, in addition to the pipeline above.
Mobile SDK engineers Six months for iOS and Android. One engineer if your app ships a single platform. 2 1 $180,000
Integrations engineer Three months of setup, a week per provider, then upkeep as their APIs change. 1 0.5 $90,000
Full build, with collection 8 $1,488,033

Not included, and not estimated here: ethics board administration and protocol review; participant recruitment and incentives across the collection window; ground-truth instrumentation for validation; cloud infrastructure during development, before any optimization; architectural rework across the collection and integration layers, which took us more than one pass. Rate benchmarks: Built In: ML engineer salary, US , Motion Recruitment: ML salary guide 2026 .

And then it has to run

Building it once is a project. Running it is permanent, and it scales with every user you add. Across four generations of our own pipeline, run-rate fell as tooling and architecture changed. Those were changes we could only make once we knew what the models needed.

Rebuild, streaming
10x
Rebuild, retooled
4x
Current pipeline
1x

Cost to run, relative to the current architecture. A first-generation pipeline is not the efficient one.

Money buys people. It does not buy elapsed iteration.

You can hire two researchers tomorrow. They cannot start until a year of collection has finished. You can hire a data engineer tomorrow, but the architecture decisions are not purely engineering ones, which is why we rebuilt the pipeline three to four times. Those were not failures. They were the architecture catching up to what the research turned out to require.

And fifteen months is longer than the window you are deciding in. Longer than most runways between raises, longer than a roadmap horizon, longer than many people stay on a product. A team choosing a vendor this quarter wants to ship this year. Start the build today and it ships in November 2027.

Questions

What buyers ask at this point

A competitor quoted us less. Why is Sahha more expensive?

Check what the quote covers first. Raw data access and a derived analytics layer are priced separately by most providers, and a scoring add-on can roughly double an entry-level quote. A quote that excludes scoring is not a quote for the same product, so the comparison has not happened yet. Once both cover the same scope, the gap is usually much smaller than it first appears. The provider-by-provider detail is on our comparison pages.

What if we take the cheaper feed and build the analytics layer ourselves?

That is a real option, and for some teams it is the right one. On the fastest defensible path it is about 15 months and 6.5 person-years of specialist research and engineering, and the first 8 of those months are ethics approval and data collection, before a model can be trained at all. We took longer. The honest question is not whether your team could do it. It is whether a validated analytics layer that lands fifteen months after you start is any use to the product you are shipping this year.

Could we start from open source instead?

Open source changes your acquisition cost to zero and leaves the engineering, validation and running costs intact. Beiwe, mindLAMP, RADAR-base and AWARE are genuinely good research platforms, and if you are running a study rather than shipping a product they may be exactly right. They do not remove the need for collection, validation, or an architecture that holds cost down at production volume. They are a way of building, not an alternative to it.

How much does it cost to run once it is built?

More than you will model, for longer than you expect. Health data is high-frequency and wants to be handled as a stream. Across four generations of our own pipeline, earlier architectures cost roughly 10x and then 4x more to run than the current one. Those reductions came from tooling and architecture changes we could only make after learning what the models actually needed. A first-generation pipeline is not the efficient one, and the run-rate scales with every user you add.

Can we not just hire two data scientists?

You can hire the people. What you cannot hire is the sequence. Researchers cannot begin until collection finishes, so a team hired on day one spends its first year unable to start. And the pipeline architecture is not purely an engineering decision. What to store, at what resolution, over what windows, with what baselines, all depend on knowing what the models will need. Staffing an engineer first and a scientist second is the ordering that produces a technically excellent pipeline built for the wrong requirements.

Bring us your comparison.

Send the quote you're evaluating and we'll tell you where another provider fits better.

30-day free trial · No credit card