The cheaper quote moves the cost to your roadmap.
Health data platforms sell two things. First the raw feed from phones and wearables. Then the analytics layer above it: scores, biomarkers, archetypes, tags and insights. Most price separately, so a cheaper quote is often not a quote for the same product.
Sahha is a health data API that includes both. This page sets out what the second layer costs to build, using our own numbers, so you can compare quotes at equal scope.
- to build the analytics layer
- 15 months
- before modelling can start
- 8 months
- for the analytics layer
- $1.2M in salary
- a first-generation pipeline
- 10x to run
Check what the quote covers.
Is scoring included, or an add-on?
Other health APIs price raw access and derived analytics separately. A scoring add-on can roughly double an entry-level quote, so a cheaper headline number may not cover what you are buying the platform for.
Sahha Included at every paid tier. Scores, biomarkers, archetypes, tags and insights ship together.
What happens when a user has no wearable?
Scores built on physiological inputs return null without them. A quote that assumes wearable coverage is a quote for the subset of your users who own one.
Sahha Every score works phone-only. A wearable adds signal, it is never required.
How long before a new user sees anything?
Baseline requirements of 7, 14, 28 or 30 days are common before a score computes at all. That window is dead time in your onboarding, every time.
Sahha Immediate in almost all cases, and calibrated from day one, because backfill supplies the history.
Who owns validation, and the liability?
Customizable scoring sounds like flexibility. In health it means the validation, and the liability for an unvalidated number someone changes their life on, transfers to you.
Sahha Validated against our own collected data, by researchers who publish it.
Provider-by-provider detail is in the comparisons. What is included at each tier is on pricing.
Fifteen months, on the fastest path there is.
Critical path
15 months8 months before a model can be trained. Headcount does not shorten it.
Running alongside
Four generations, 24 months| Phase | Duration | Budget shortens it? |
|---|---|---|
| Ethics protocol A commercial review board, which is the fast option. Ours took closer to four through an academic one. | 2 months | No |
| Longitudinal data collection The aggressive floor, and it costs you a full seasonal cycle. We collected for nine to twelve. No number of hires shortens either. | 6 months | No |
| Pipeline, across 3 to 4 rebuilds Each rebuild took roughly six months. No first plan contains four of them. | runs alongside | Partly |
| Discarded model approaches The approaches that do not work, before the one that does. Ours took three to four months. | 2 months | Partly |
| Scores Two to three weeks each once the data and the approach exist. Readiness took a month. | ~3.5 months | Yes |
| Archetypes Built on the same foundation. | 1.5 months | Yes |
| Insights, trends and comparisons The easy part, and the part most teams imagine is the whole job. | ~2 weeks | Yes |
Fifteen months only holds if you take all four of these. We took longer on each.
- Commercial IRB instead of an academic board
- Realistic, and genuinely faster. Costs money rather than time, and needs someone who can write the protocol, which means knowing what you intend to collect and why.
- Six months of collection instead of nine to twelve
- Captures half an annual cycle. Sleep, activity and mood all move with season and daylight, so a single-season dataset carries a bias you will not detect until you deploy into the other half of the year.
- Fewer repeated clinical measures per participant
- Thinner validation, most visibly for anything mental-health related, where the label is the expensive part.
- Validate fewer scores
- The fastest real saving. It also means shipping less than you set out to build.
What it costs in people
These are our own figures, from building this layer ourselves: ethics protocol, longitudinal collection, models, and four generations of pipeline. Scoped to what you would replicate, so front-end and dashboard work is excluded. Rates are published benchmarks, not our payroll, so the model stays checkable.
| Role | FTE | Person-years | Cost |
|---|---|---|---|
| ML researchers (PhD) Modelling cannot begin until collection completes. Costed as hired late, the cheaper of the two options. | 2 | 1.5 | $318,033 |
| Pipeline engineers | 2 | 4 | $720,000 |
| Senior infrastructure Part allocation. The rest goes to API, webhooks and dashboard work you would not replicate. | 0.5 | 1 | $180,000 |
| Total | 6.5 | $1,218,033 | |
| Mobile SDK engineers Six months for iOS and Android. One engineer if your app ships a single platform. | 2 | 1 | $180,000 |
| Integrations engineer Three months of setup, a week per provider, then upkeep as their APIs change. | 1 | 0.5 | $90,000 |
| Full build, with collection | 8 | $1,488,033 | |
Not included, and not estimated here: ethics board administration and protocol review; participant recruitment and incentives across the collection window; ground-truth instrumentation for validation; cloud infrastructure during development, before any optimization; architectural rework across the collection and integration layers, which took us more than one pass. Rate benchmarks: Built In: ML engineer salary, US , Motion Recruitment: ML salary guide 2026 .
And then it has to run
Building it once is a project. Running it is permanent, and it scales with every user you add. Across four generations of our own pipeline, run-rate fell as tooling and architecture changed. Those were changes we could only make once we knew what the models needed.
Cost to run, relative to the current architecture. A first-generation pipeline is not the efficient one.
Money buys people. It does not buy elapsed iteration.
You can hire two researchers tomorrow. They cannot start until a year of collection has finished. You can hire a data engineer tomorrow, but the architecture decisions are not purely engineering ones, which is why we rebuilt the pipeline three to four times. Those were not failures. They were the architecture catching up to what the research turned out to require.
And fifteen months is longer than the window you are deciding in. Longer than most runways between raises, longer than a roadmap horizon, longer than many people stay on a product. A team choosing a vendor this quarter wants to ship this year. Start the build today and it ships in November 2027.
What buyers ask at this point
A competitor quoted us less. Why is Sahha more expensive?
Check what the quote covers first. Raw data access and a derived analytics layer are priced separately by most providers, and a scoring add-on can roughly double an entry-level quote. A quote that excludes scoring is not a quote for the same product, so the comparison has not happened yet. Once both cover the same scope, the gap is usually much smaller than it first appears. The provider-by-provider detail is on our comparison pages.
What if we take the cheaper feed and build the analytics layer ourselves?
That is a real option, and for some teams it is the right one. On the fastest defensible path it is about 15 months and 6.5 person-years of specialist research and engineering, and the first 8 of those months are ethics approval and data collection, before a model can be trained at all. We took longer. The honest question is not whether your team could do it. It is whether a validated analytics layer that lands fifteen months after you start is any use to the product you are shipping this year.
Could we start from open source instead?
Open source changes your acquisition cost to zero and leaves the engineering, validation and running costs intact. Beiwe, mindLAMP, RADAR-base and AWARE are genuinely good research platforms, and if you are running a study rather than shipping a product they may be exactly right. They do not remove the need for collection, validation, or an architecture that holds cost down at production volume. They are a way of building, not an alternative to it.
How much does it cost to run once it is built?
More than you will model, for longer than you expect. Health data is high-frequency and wants to be handled as a stream. Across four generations of our own pipeline, earlier architectures cost roughly 10x and then 4x more to run than the current one. Those reductions came from tooling and architecture changes we could only make after learning what the models actually needed. A first-generation pipeline is not the efficient one, and the run-rate scales with every user you add.
Can we not just hire two data scientists?
You can hire the people. What you cannot hire is the sequence. Researchers cannot begin until collection finishes, so a team hired on day one spends its first year unable to start. And the pipeline architecture is not purely an engineering decision. What to store, at what resolution, over what windows, with what baselines, all depend on knowing what the models will need. Staffing an engineer first and a scientist second is the ordering that produces a technically excellent pipeline built for the wrong requirements.
Bring us your comparison.
Send the quote you're evaluating and we'll tell you where another provider fits better.
30-day free trial · No credit card