Skip to main content

What a health-app prediction can — and cannot — tell you

A confident-looking score is only the beginning. Three questions to help you understand health-app predictions, evidence and uncertainty.

· By Alexandra Mont · 5 min read

You have plans for tomorrow. Then an app suggests it could be a difficult day.

Would you change your plans? Ignore the prediction? Look for an explanation?

Before a number becomes part of that decision, I want to understand what it means. What is the app estimating? How was it tested? What happens when it does not have enough information?

These are questions I ask as the founder of ENDOless. They belong on our side of the screen, too.

Health technology should make it easier to understand our experiences. For me, that includes being able to question the result.

1. What is actually being predicted?

“Predicts your health” leaves a lot unanswered.

Tomorrow’s symptom score? A date range? The likelihood of a particular event? Each claim needs an explanation: what outcome, over what period, and to help you do what?

It also helps to separate the different things an app can show you:

  • A record describes information entered or collected: the symptom you logged yesterday.
  • A summary brings records together: your average reported symptom score over a week.
  • A prediction estimates something that has not happened yet, or is not directly known. It can be wrong.

A forecast does not, by itself, establish a diagnosis. Yet records, summaries and predictions can appear in similar charts, with equally confident language.

I want the distinction to be clear before someone has to decide how much weight to give a result. Understanding the claim should not require understanding the code.

2. What evidence supports the number?

Imagine an app that says, “It won’t rain today,” every single morning.

A percentage can look convincing before you know what was counted.

Here is a fictional example. Imagine a test covering 100 days:

Imagine an app that says, “It won’t rain today,” every single morning.

Over 100 days, there are 90 dry days and 10 rainy days.

The app gets the weather right on 90 days. That sounds good — until you realise it missed every rainy day.

If you needed to know when to bring an umbrella, it would never help you.

A health prediction can have the same problem.

We need to ask: does it spot the events we care about, or mostly get the easy days right?

This is arithmetic, not a study.

The point is that an overall score can hide the very errors you care about.

If I were deciding whether a prediction was useful, I would want to ask:

  • What counted as a correct prediction?
  • How often did the system miss an event? How often did it warn about one that did not happen?
  • Was it tested on people whose data were not used to build it?
  • How similar were those people and their circumstances to its intended users?
  • Can I read the evaluation?

These questions reflect wider concerns in health AI. WHO’s regulatory considerations highlight intended use, external validation and data quality. In France, HAS’s descriptive framework for AI medical devices asks manufacturers to describe their claimed use, data and performance in the reimbursement-evaluation process.

Those frameworks have specific purposes. My practical takeaway is simpler: a provider should be able to explain what its evidence supports, and where that evidence stops.

3. What are its limits?

Perhaps you stopped recording for a few days. You were busy, tired, or wanted a break. Now the app has a gap.

A missing entry does not mean you had no symptoms.

I want a tool to distinguish between “you recorded that this was absent” and “we do not know”. I also want it to explain when it has too little information for a useful estimate.

Can it show uncertainty? Can it say, “There is not enough information today”? Or does every result look equally confident?

Limits matter when interpreting patterns, too. If poorer sleep and higher symptom scores appear together, that observation alone does not establish which caused which, or whether something else influenced both. It can be a starting point for a conversation with a healthcare professional.

A general health-app score or forecast should not be a reason to change medication or delay care. Follow a medical device’s stated instructions and discuss care decisions with an appropriate healthcare professional.

When an app’s promise goes further than its evidence

Recording symptoms can be useful. Helping someone organise their history can be useful. Neither, on its own, demonstrates that an app can predict a disease-related event or guide a care decision.

The evidence has to support the specific promise: the outcome, the people it is intended for and the way it will be used. Calling a feature “AI-powered” does not answer those questions. Neither does a high accuracy figure without an explanation of the test and the errors.

There is also a regulatory boundary. In the EU, software intended for a medical purpose can fall under medical-device rules. Those purposes include predicting or monitoring disease. General wellbeing software does not automatically qualify, and neither does every app that stores health information. The intended purpose matters, including what the manufacturer says in its marketing. MDCG software guidance, revised June 2025.

My position is that a provider should keep its promises within the evidence it can show and meet the regulatory requirements that apply to those promises. If it has demonstrated tracking, it should explain the value of tracking. A medical prediction requires its own justification.

Why we are pursuing the medical-device route

This is why we are pursuing a medical-device pathway for ENDOless’s intended medical functions. I want the claims we eventually make to be supported by evidence people can examine, with clear limits on what the product is intended to do.

That pathway brings concrete responsibilities: defining the intended use, managing risks, documenting safety and performance, and planning and maintaining a clinical evaluation. The applicable conformity-assessment procedure depends on the device’s risk class; independent notified-body involvement is required where the rules call for it. France’s G_NIUS guide to CE marking explains these steps and why the work begins during development.

For me, the value is in the work behind the claim. We need to understand what a wrong prediction could mean for someone, what evidence would support a useful result, and when the system should acknowledge that it cannot tell them enough.

Medical-device status follows the intended purpose and applicable rules; it is not an optional quality label for a product that already has a medical purpose. Nor would CE marking make every estimate correct or support claims outside the assessed scope.

We are describing the pathway we are pursuing, not announcing completed CE marking or a clinically validated ENDOless prediction feature. The exact scope and claims must follow the evidence and regulatory assessment.

The person using a health app already has enough uncertainty to manage. Our responsibility is to make clear what we know, what we are estimating and what we cannot yet support.

Before you trust the next prediction, keep these three questions:

What is being predicted?
What evidence supports it?
What are its limits?

If the answers are hard to find, ask the provider.

Including us.


Source note:

About the author

Alexandra Mont Alexandra Mont
Updated on Sep 21, 2026