Short answer: FDA cleared means the software resembles an older product enough to be sold. It does not mean anyone checked whether patients lived longer or felt better. Of 1,357 FDA cleared AI medical devices (AI, artificial intelligence, is software that learned a pattern from old cases, then applies it to a new one) authorized through December 5, 2025, only 3 were tested on a result the patient actually lives.
That number is back in the news this weekend because Earth.com walked through it on October 4, and AI news indexes picked it up on October 5. The study itself is not from this week. Treating 1,357 as today’s live total would be wrong. The useful part is the idea under the number, and that idea does not expire on Friday.

What is actually in the room with the patient?

A scan is a picture. The new product is software that comments on the picture.
Start before the label. A hospital scan is a picture of the inside of a body. A CT (computed tomography, many X-ray slices stacked into a 3D picture) and a mammogram (an X-ray of the breast, used to look for cancer) are just pictures plus measurements. A radiologist (a doctor whose job is reading those pictures) looks at them and says what they think is there.
The thing the FDA (the U.S. Food and Drug Administration, the agency that decides which medical products may be sold in the United States) has been clearing is usually not a robot and not a chatbot. It is a medical device (any product meant to diagnose or to guide treatment, and software counts) that was built with machine learning (finding a pattern from examples, instead of a person writing every rule by hand). Artificial intelligence (software that learned a pattern, then applies it to a new case) is the broad name. The device is the legal object.
In the room, that software might circle a spot on a lung, rank which scan a doctor should open first, or flag a possible stroke. Those are assistive jobs. The patient still meets a person. The badge on the software can still be misunderstood, because “allowed to be sold” and “shown to help” are different sentences.

What does cleared actually check?
It checks resemblance to an older product. A new medicine has to clear a higher bar.
Most of these tools use the 510(k) (a shortcut application, named after a section of U.S. law, that lets a maker sell a device by comparing it with one already sold). The older product is the predicate device (the product you claim to resemble). If the FDA agrees you are substantially equivalent (close enough in job and in technology that you are not treated as brand new), you get clearance (permission to sell because of that resemblance).

Hold the three doors apart, because headlines mash them into “FDA approved.” Approval (the stricter yes, used when the product itself must be shown safe and effective) is not what most of this software received. PMA (premarket approval, the door for the highest-risk devices, which demands clinical proof) is rare here. De Novo (the door for a genuinely new lower-risk device that has no older twin, which still requires the maker to show its own safety and effectiveness) is also not the usual path. The paper says the vast majority reached the market through 510(k), where an independent prospective trial (a human study whose question is locked before anyone is enrolled) is not required.
A new medicine does not get the resemblance shortcut. It has to show, in people, that it does some good. Sebastián Andrés Cajas Ordóñez, one of the study authors, told Earth.com on October 4, 2026: “FDA clearance means the device resembles one already on the market. It does not mean it helps anyone.”
How does 1,357 become 3?
The team listed every authorized AI device, then asked which ones had a public study of patient benefit.
The paper is “1,357 AI medical devices cleared, 3 actually tested on patient outcomes,” in PLOS Digital Health (an open-access journal, meaning anyone can read the paper without paying), August 19, 2026. Lead author Rawan Abulibdeh is at the University of Toronto. Coauthors include researchers at MIT Critical Data (a Massachusetts Institute of Technology group that studies hospital data). The DOI (a permanent ID for the paper) is 10.1371/journal.pdig.0001597.
They matched the FDA’s public AI-device list, frozen at December 5, 2025, to ClinicalTrials.gov (the U.S. site where planned human studies are supposed to be registered before they start) and to PubMed (the public index of medical papers). They were not grading secret company slides. They were grading what a doctor, a journalist, or a patient can actually look up.
Here is the funnel, in counts, not in adjectives.
| Stage | Devices | Share of 1,357 | What that stage means |
|---|---|---|---|
| Authorized by the FDA through Dec 5, 2025 | 1,357 | 100% | Allowed to be marketed |
| Linked to a registered prospective trial | 34 | 2.5% | Someone planned a human study in public |
| Trial results posted | 12 | 0.9% | Numbers from that study were put on the record |
| Peer-reviewed paper | 12 | 0.9% | Outside scientists checked a write-up before publication |
| Tested on a patient-centered outcome | 3 | 0.2% | The study asked whether patients did better |
A patient-centered outcome (a result the patient lives, rather than a score the software gets) means mortality (how often people die), morbidity (how often people are seriously harmed or get sicker), or readmission (coming back into the hospital soon after going home). Quality of life sits in the same family. The paper’s public summary does not name the 3 devices.
The 12 with posted results and the 12 with a peer-reviewed paper (a paper other scientists checked before a journal printed it) are both 0.9%. The paper does not say they are the same 12. Count them as two separate rungs unless you open the supplement.
Why is “the AI was accurate” a different question?
Accuracy asks whether the software agreed with a label someone already made. Benefit asks whether the person ended up better.
Accuracy (how often the software’s answer matches a diagnosis already written down, or a mark an expert already drew) can be excellent while nothing about the patient’s life changes. It can also be excellent while the patient does worse, if people trust the circle and stop looking.

Picture a tool that circles every shadow a senior radiologist would also circle. Agreement is high. If that radiologist was already going to see the scan, the patient may not live longer. If the circle makes the next reader hurry, the patient can do worse. Neither ending is inside an accuracy percentage.
The authors put the consequence in one line: readiness should no longer be defined by FDA clearance alone, but by demonstrated, durable, and equitable benefit. Their shorter version is the one worth keeping: “AI tools must be life-tested before they can be called life-saving.”
That is not an argument for ripping software out of reading rooms. A triage tool (software that only decides which scan a human opens first) might still shorten the wait for a stroke read. The honest sentence is “we have not shown the patient outcome,” not “the tool does nothing.”
Where did almost all of these devices land?
In the specialty that reads pictures, which is also the specialty with the thinnest public trial record.
Of the 1,357 devices, 1,059 are radiology (the specialty that reads scans such as CT and MRI), which is 78%. Heart and blood-vessel software is about 9% of the list. Brain and nerve software is about 5%. Everything else is about 8%. Anesthesiology (the specialty that keeps a person stable during a procedure) had 22 cleared devices and zero registered prospective trials. A broader “other” bucket of 88 devices had the highest trial rate the authors report, 14.8%, which is still a small minority.
| Specialty | What the paper says | Registered prospective trials |
|---|---|---|
| Radiology | 1,059 devices, 78% of the list | Under 1%. Write-ups of the paper put it at 3 of 1,059 |
| Cardiovascular | About 9% of devices | About 9.5% of those devices |
| Neurology | About 5% of devices | About 9.7% of those devices |
| Anesthesiology | 22 devices | None |
| Other | 88 devices in the authors’ other bucket | 14.8% |
Read the two threes separately. Three devices in the whole list were tested on patient-centered outcomes. Three radiology devices had a registered prospective trial. The paper does not say those are the same products. Merging them is how a true statistic becomes a false one.
Asked which group worried him once he saw the evidence, Cajas Ordóñez said radiology, “because of the scale.” One thousand products, three public trials. The scale is the point.

Who was missing from the few studies that exist?
The studies are small, mostly American, mostly paid for by the seller, and they often skip the people who are hardest to treat.
Among the 34 registered trials, 73% enrolled fewer than 500 people, and about one quarter enrolled fewer than 100. Sixty-eight percent ran only in the United States. Thirty-two of the 34 (94%) were industry-led (designed and paid for by the company that sells the device). Nine of the 34 reported any subgroup analysis (a second look that asks whether the result still holds for a smaller group, such as older adults). Of those, 5 looked at sex, 4 at age, 3 at race or ethnicity, and none at language.
Pregnancy was an exclusion in 42% of the cardiovascular trials and 33% of the radiology trials. Children were almost universally left out. Non-English speakers show up on the exclusion figure. Fifty-nine percent of the trials were about diagnosis (naming a disease that may already be there), 21% about screening (checking people who feel well, hoping to catch disease early), and 9% about guiding treatment.
Most of the evidence the team could find, 62%, was observational (researchers watched records that already existed, instead of assigning some patients to the tool and some patients away from it). An observational study can suggest a pattern. It cannot, by itself, prove the tool caused a better outcome.
Cajas Ordóñez also said the quiet commercial part out loud: devices validated on one American patient population are then sold worldwide. A passing score on one hospital’s archive is not a passport that the next hospital’s patients match the archive.

What would make this post false?
A newer census, a hidden public trial, or a reader who treats “not proven” as “proven useless.”
The authors say their method can undercount proprietary validation studies (private tests a company ran and never published). A good study locked in a filing cabinet would not appear. A bad one would not appear either. Clearance would still not equal a public proof of benefit. They also did not build a non-AI comparison, did not pull every recall report, and did not score how much the software acts alone versus how much a doctor must confirm it.
On October 4, 2026, Cajas Ordóñez told Earth.com the team has not repeated the census. He would not claim that trial registration has improved, because he had not counted again. The FDA list kept growing after December 5, 2025. If you need today’s headcount, download the FDA AI-enabled device list and do not quote 1,357 as if it were October 5.
Two later regulatory facts belong next to the paper, not instead of it. On September 17, 2026, the FDA published a final order refusing to excuse four families of radiology computer-aid software from a fresh 510(k). The resemblance door stayed a requirement. That order does not add the patient-outcome test this paper is describing. Separately, the FDA’s device center has a discussion paper on generative AI (models that write new output, not only circle a known pattern) medical devices. Public comments on docket FDA-2026-N-7874 are due October 19, 2026. The questions in that paper, including what happens when a model confabulates (states a fluent falsehood), are the next argument. They do not fill in the missing outcome trials.

What do you ask before you trust the badge?
Three questions. They take less time than reading the press release.
Cajas Ordóñez offered the practical version. The definitions sit beside each question so you can use them on a pitch, a hospital brochure, or a model card (the short document that says what a model was trained to do).
- Was this tested on patients like me, or only on stored images and records? A stored-image test never sees what happens after the doctor acts.
- What did the test measure: accuracy, or whether patients actually did better? If the endpoint (the result the study promised to count) is agreement with an old label, it is not a benefit study.
- Who was left out? If pregnancy, age over 75, children, or a language you speak was an exclusion, the average is not your average.

If you build software and you want the sentence “cleared” to mean more than resemblance, the missing work is a prospective study with a patient-centered endpoint, registered before you enroll anyone, large enough to include the people you plan to sell to. Clearance can still be the legal door. It is just not the scientific one.
More on AI landing in places where mistakes reach people: an OpenAI agent that wrote files onto a Medicare server, companies that replaced humans with AI and had to reverse it, and what the Super Intelligence accord does and does not enforce.
Common questions about FDA cleared AI medical devices
Does FDA cleared mean FDA approved?
No. Cleared, for these devices, almost always means a 510(k) resemblance decision. Approved usually means PMA, the higher door with clinical proof. News copy swaps the words. The paperwork does not.
Were only 3 AI devices ever studied at all?
No. Thirty-four had a registered prospective trial, and many more have accuracy tests on stored scans. Three had been studied for mortality, morbidity, or readmission. “Barely tested on patient outcomes” is the claim. “Never tested on anything” is a different claim, and it is false.
Are the 3 outcome devices the same as the 3 radiology trials?
The paper does not say that. Keep them as two different counts.
Is 1,357 still the size of the list?
No. That is the paper’s cutoff on December 5, 2025. The authors said on October 4, 2026 that they had not counted again.
Should a hospital turn these tools off?
The paper does not say that. It says clearance is not evidence of benefit, and that benefit should be measured, including for people the first studies left out.
Clearance tells you the product resembles something already sold. It does not tell you the patient did better.





1 comment