AI for Medical Devices: A Practical Guide for 2026

AI for Medical Devices: A Practical Guide for 2026

2026-08-10 · Tommaso Maria Ricci

Regulators have now authorized more than fifteen hundred AI-enabled medical devices in the United States alone. The FDA's own list of AI-enabled medical devices carried 1,524 entries as of its 30 March 2026 update, with 1,164 of them, roughly three quarters, sitting in radiology. The 2025 cohort alone accounts for 333 authorizations. AI for medical devices stopped being a research topic somewhere around 2023. It is now a regulatory workload.

That distinction matters more than any capability claim you will read this year. Most of the money a device company will spend on AI over the next thirty-six months has nothing to do with putting a model inside the product. It goes into the paperwork, the evidence, the post-market surveillance and the quality system that surround the product.

This guide separates the two. The first half covers where AI produces measurable returns inside a medical device business, in functions that have nothing to do with the device itself. The second half covers what happens when the AI is in the product, which is where the regulatory picture changed twice in the last eighteen months and where most published advice is now out of date.

One methodological note. You will not find a list of tools here. Tools change every quarter and picking one is the least important decision in the entire sequence. You will find the criteria for deciding where to apply AI, how to measure whether it worked, and how to avoid the mistake that kills roughly a third of these projects before they reach production.

What actually changed: AI in the product versus AI in the company

The confusion between these two things produces most of the wasted budget in this sector, so it is worth being precise.

AI in the product means a model that contributes to the device's intended medical purpose: triage, detection, measurement, dosing support, signal interpretation. This is a regulated design decision. It affects your classification, your clinical evidence, your technical documentation, your notified body conversation and, in the European Union, a second regulation stacked on top of the first.

AI in the company means everything else: regulatory writing, literature review, complaint triage, supplier documentation, sales enablement, field service. None of this touches the device. All of it is where the near-term return lives.

The failure mode is predictable. A device company decides to "do AI", the conversation immediately becomes a product conversation, the product conversation immediately becomes a regulatory conversation, and eighteen months later there is a roadmap and no result. Meanwhile the regulatory affairs team is still assembling technical files by hand.

The companies getting value are doing the reverse. They start where the work is textual, repetitive and internally reversible, they build the internal capability there, and they carry that capability into product decisions once they have earned the credibility. The broader version of this argument is in the guide on enterprise AI adoption.

Where AI pays off inside a medical device company

The areas below are ordered by the ratio of implementation difficulty to measurable impact. Not by how impressive they sound in a board deck.

1. Regulatory documentation and technical file assembly

This is the highest-return, lowest-risk starting point in the entire sector, and it is the one most companies skip because it feels unglamorous.

A device company generates an enormous volume of structured text that must be internally consistent: technical documentation, design history records, risk management files, clinical evaluation reports, post-market surveillance plans, periodic safety update reports, declarations of conformity. Much of it restates the same underlying facts in different formats for different audiences.

A model that reads the source records and drafts the document in the required template removes the slowest and least skilled part of the work. The regulatory professional moves from writing to reviewing, and the same team handles a materially larger portfolio.

The measurement is straightforward because before and after are comparable: how many hours did that document take, how many does it take now. No attribution model required.

The one non-negotiable rule is that a named person reviews and signs. Not for formality, but because regulatory responsibility stays with a human being, and an error in a submitted document is a legal problem rather than a productivity problem.

2. Literature review and clinical evidence work

Clinical evaluation under the European framework requires systematic literature review, and it requires it repeatedly rather than once. The same is true of state-of-the-art justification, competitor equivalence arguments and post-market clinical follow-up.

The work is genuinely well suited to language models: large volumes of text, defined inclusion and exclusion criteria, structured extraction into evidence tables. A system that screens abstracts against your criteria and drafts extraction tables with citations compresses weeks into days.

The condition, and it is absolute, is traceability. Every extracted claim must point back to the source document and the specific passage. A system that summarizes without citing is unusable here, because the entire artifact exists to be audited by someone who will check.

Used correctly this is also a quality improvement rather than only a speed improvement, because consistent criteria applied by a system beat inconsistent criteria applied by three people under deadline pressure.

3. Complaint handling and post-market surveillance

This is the area where the operational pain is highest and where the regulatory clock is least forgiving.

Complaints arrive as unstructured text from distributors, clinicians, service engineers and patients, in multiple languages, in formats that resist automation. Someone has to read each one, decide whether it is reportable, classify the failure mode, link it to the right product family and start the clock.

A model that reads the incoming text, extracts the structured fields, proposes a classification and flags the ones that look reportable turns a reading task into a verification task. Throughput goes up and, more importantly, the variance in classification between reviewers goes down.

Two safeguards decide whether this works. The first is a confidence threshold below which the item routes to a human queue instead of being auto-classified. The second is a fixed control sample re-checked manually every week even when everything looks fine, because these systems degrade quietly when incoming formats change.

Reportability decisions themselves stay human. The system narrows and prepares; it does not decide.

4. Quality system operations and supplier management

Every device manufacturer maintains an archive that nobody can practically consult: supplier agreements, validation records, change control history, audit findings, corrective action files, calibration records.

A system that answers questions against that archive with citations turns a dead archive into a usable resource. The useful questions are mundane and valuable: which suppliers have open corrective actions, which validation records are due for review this quarter, what did we conclude about this failure mode three years ago, which agreements contain the clause the auditor just asked about.

Audit preparation is the obvious application, and the one that pays for the project on its own in most mid-sized companies. Preparing for a notified body audit is largely an information retrieval exercise performed under time pressure by people who have other jobs.

The same traceability rule applies. In a regulated archive, an answer without a citation is worse than no answer, because it invites a claim nobody can substantiate.

5. Manufacturing quality and inspection support

For companies that manufacture rather than outsource, visual inspection and process monitoring are the classic applications, and they predate generative AI by a decade.

What has changed is the cost of entry. Building an inspection model used to require a data science team and months of labeled images. It now sits within reach of a mid-sized manufacturer with a well-defined defect taxonomy and a competent engineering team.

The constraint that catches people is validation. In a regulated production environment an inspection system is part of the process, which means it needs qualification, change control and a defined behavior when it fails. Companies that treat it as an IT project rather than a process change discover this during their next audit.

The broader operational framing is in the guide on AI for manufacturing.

6. Commercial operations, tenders and reimbursement support

Medical device selling is document-heavy in a way that few other B2B sectors are: hospital tenders, framework agreements, reimbursement dossiers, clinical value summaries adapted per country and per payer.

The pattern that works is variation on a verified base. Take one approved clinical value argument and produce twenty versions for different tenders, specialties and reimbursement environments. The model is excellent at varying and poor at inventing the underlying claim. That distinction decides whether the project produces value or regulatory exposure.

The pattern that does not work is generating clinical claims. In this sector a claim outside the approved intended use is not a marketing mistake, it is a compliance event. The operating rule I use is simple: if a statement is not already supported in an approved document, it does not leave the building.

7. Field service and technical support

Devices in the field generate service calls, and service organizations carry a knowledge distribution problem: the best engineer knows things the newest engineer does not.

A retrieval system over service manuals, historical tickets and known issues, surfacing suggested diagnoses to the engineer rather than to the customer, compresses the learning curve. The evidence for this pattern is strong and is worth citing rather than asserting.

The study Generative AI at Work by Brynjolfsson, Li and Raymond, conducted on 5,179 customer support agents, measured a 14% average productivity increase, with a 34% improvement for the least experienced workers and minimal effect on the most experienced ones. The mechanism the authors identify is as interesting as the result: the system propagates the practices of the strongest performers and shortens the ramp for new hires.

Applied to a device service organization, that means the value is not making your senior engineer faster. It is bringing a new hire to competence in weeks instead of years, in an environment where competence has safety implications.

AI inside the product: what regulators actually require in 2026

This is where most published advice is out of date, because the framework changed twice in eighteen months.

The United States: authorizations and change control

The FDA has been authorizing AI-enabled devices for years through the existing pathways, and the volume has accelerated sharply. The practical implication is that AI in a device is not novel from a regulatory standpoint. It is a well-trodden path with known evidence expectations.

The change that matters most operationally is the Predetermined Change Control Plan. On 3 December 2024 the FDA published final guidance titled Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions.

A PCCP lets a manufacturer describe planned modifications, the methodology to develop and validate them, and an assessment of their impact, and have that plan reviewed as part of the original submission. Modifications that fall inside the approved plan can then be implemented without a new marketing submission. Notably, the final guidance applies to all AI-enabled device software functions rather than only to machine learning systems that continue to train, which was the narrower scope of the draft.

The strategic consequence is easy to miss and expensive to miss. If your product roadmap involves retraining or threshold adjustment, the PCCP is the difference between a product you can improve and a product frozen at the state it was authorized in. That decision belongs in the submission strategy, not in a later engineering discussion.

The European Union: two regulations, one device

In Europe an AI-enabled medical device sits under both the medical device framework and the AI Act simultaneously. This produced enough confusion that the regulators addressed it directly.

On 19 June 2025 the Medical Device Coordination Group and the European Artificial Intelligence Board jointly published MDCG 2025-6, a FAQ on the interplay between the MDR, the IVDR and the AI Act. It was the first joint guidance from those two bodies, and it exists because manufacturers could not tell which requirements stacked and which were satisfied by existing quality system work.

The practical reading is that requirements integrate rather than duplicate. Bias mitigation, transparency, data governance, human oversight and cybersecurity expectations from the AI Act map onto processes a compliant manufacturer already runs under the device framework. The work is extension and documentation, not a parallel system.

The deadline that moved, and the one that did not

The timing changed in 2026, and getting this wrong in either direction is costly.

The AI Act has been in force since 1 August 2024 and applies in phases. Obligations for general-purpose AI models have applied since 2 August 2025, and most of the regulation became applicable in August 2026.

Then the package known as the Digital Omnibus postponed the high-risk obligations. Standalone high-risk systems listed in Annex III moved to 2 December 2027. AI embedded in already-regulated products under Annex I, which is the category that covers medical devices, moved to 2 August 2028. The Council confirmed the agreement on 29 June 2026. The detailed reconstruction is in Gibson Dunn's analysis of the Omnibus agreement, and the official framework page is maintained by the European Commission.

Two readings follow, and both are practical. You have more time than the 2026 planning assumption most device companies were working from, which means the correct response is to use it rather than to relax. And the transparency obligations were not postponed, so they apply now to any generative system your organization uses with customers or clinicians.

There is a trap in the extension worth naming. A device on a five-year development cycle that reaches the market in 2028 is designed in 2026. Design decisions taken now on data provenance, dataset documentation and logging determine whether the 2028 conformity work is a documentation exercise or a redesign. Companies treating the delay as permission to wait are choosing the expensive path without noticing.

The general compliance framing is in the guide on AI for compliance, and the organizational side is in the guide on AI governance for business.

What does not work, and nobody says in a sales deck

The credibility of a guide is measured in this section.

The general-purpose assistant produces no measurable result. Giving everyone access to a chatbot and hoping something emerges is the most common adoption mode and the lowest-return one. It produces a diffuse, invisible improvement that appears in no financial statement. It is not useless, but it is not a project, and it is why many leadership teams believe they have adopted AI without having changed anything.

AI does not compensate for a broken process, it accelerates it. If your document approval path has four unnecessary steps, generating documents faster means waiting longer in the same four steps. Process redesign comes first.

It does not replace domain competence, it amplifies it. A model that drafts a clinical evaluation report drafts it well only if the reviewer knows what is missing. In the hands of someone who does not, it produces professional-looking documents with wrong content, which is the worst possible combination in a regulated environment.

Hallucinations are not solved by better prompting. They are reduced by grounding in your documents and mandatory source citation, and managed by human verification. Anyone claiming their system does not hallucinate is selling.

Catalogue training does not produce adoption. Generic AI courses have a very low transfer rate to actual work. What works is training on the person's own process, using their real documents, in an hour, with an output they keep.

The aggregate picture supports the caution. According to the Gartner forecast cited in research published in MIT Sloan Management Review, roughly 30% of generative AI initiatives were expected to be abandoned after proof of concept by the end of 2025. The authors read the cause bluntly: the problem is less about technology than about knowledge, because organizational data stays fragmented, inaccessible and underused by the people making decisions.

What it costs in 2026 and how the return is calculated

Entry cost has collapsed, and that is the fact that changes the arithmetic for smaller manufacturers. According to the Stanford HAI AI Index 2025, inference cost for a system performing at the level of GPT 3.5 fell more than 280 times between November 2022 and October 2024.

The practical consequence is that the dominant cost line is no longer the model. It is integration with existing systems, getting data into usable shape, and the time of people who have to change habits.

Here are realistic orders of magnitude for a device company between 50 and 500 employees, over the first twelve months.

| Initiative | First-year investment | Expected return | Time to effect |

|---|---|---|---|

| Regulatory document drafting support | 15,000 - 60,000 USD | 30 - 50% of drafting time recovered | 45 - 90 days |

| Literature review and evidence extraction | 25,000 - 90,000 USD | Weeks compressed to days per review | 60 - 120 days |

| Complaint triage and classification | 30,000 - 120,000 USD | Higher throughput, lower classification variance | 90 - 150 days |

| Quality archive retrieval and audit prep | 25,000 - 100,000 USD | Audit preparation time sharply reduced | 60 - 150 days |

| Tender and reimbursement content variation | 15,000 - 60,000 USD | More market-specific versions per person | 45 - 90 days |

| Field service knowledge support | 30,000 - 110,000 USD | Shorter ramp for new engineers | 90 - 180 days |

| AI inside the device | 250,000 USD and up | Product differentiation, new claims | 18 - 48 months |

The ranges are wide on purpose. The variable that determines where you land is not the vendor, it is how accessible your data is. A company whose quality system exposes structured records pays roughly half what a company extracting information from scanned PDFs and shared folders pays.

That last row is deliberately separated from the others. Putting AI in the device is a product investment measured in years, evaluated against a product business case, and it should never be funded out of an efficiency budget. Companies that blur that line end up with an efficiency program that delivers nothing and a product program with no runway.

Calculating return without telling yourself stories

Return on these projects has four components with very different levels of certainty, and they should be kept separate.

The first is time recovered on a specific activity, and it is the most solid: measured before and after on a sample, with a stopwatch. The second is volume handled at constant headcount, which shows up directly in operational numbers. The third is error reduction, real but slow to demonstrate because it needs long observation windows. The fourth is quality perceived by customers and auditors, which exists but is not isolable.

A project justified only by the third and fourth components is not justified. A project that pays for itself on the first two, with the others arriving as a bonus, is solid. The full calculation method is in the guide on AI ROI for business.

Data, confidentiality and validated systems

This section is worth more than the rest combined, because it is where device companies get hurt.

Three checks before connecting any generative system to company data.

Where the data is processed. The processing jurisdiction belongs in the contract. A vendor who answers this vaguely is describing a risk, not a service.

Whether inputs are retained, and for how long. Many services retain inputs for defined periods for security purposes. It is a contractual term, not a technical detail, and it has to be weighed against the nature of what you intend to send.

Whether your data trains third-party models. There is one acceptable answer, and it is no, written into the contract rather than stated on a call.

Device companies carry a fourth question that other sectors do not. If a system participates in a process governed by your quality system, it needs to be handled as a validated system: intended use documented, validation evidence retained, change control applied, and a defined behavior when it fails or is unavailable.

The failure mode I see most often is a team quietly adopting a general-purpose assistant for complaint triage or document drafting, producing real value for six months, and then discovering during an audit that a quality-relevant process has been running on an unvalidated, undocumented tool. Retrofitting that documentation costs more than doing it correctly at the start, and the finding is avoidable.

The practical rule: decide early whether a given use is quality-relevant. If it is, it enters the quality system on day one. If it is not, say so explicitly and write down why.

Real cases: what happens when it works

The projects I have run over the last several years span very different industries, and the pattern repeats with a regularity that stopped surprising me a while ago.

WSB Sport, 30% sales increase. Marketing reorganization with AI support for segmentation and content production. The transferable point for device companies is that the gain did not come from volume of content produced, it came from the ability to differentiate the message per segment without adding people. That is exactly the tender and reimbursement problem.

Hospitality business, revenue from 9 to 10 million. Work on demand forecasting and allocation. The interesting detail is that the necessary data had been sitting in the management system for years and had simply never been read together. That is the most common condition I encounter, and it describes most quality archives.

Medical center, 20% capacity increase. No new equipment, no new hires. Better slot allocation based on real duration analysis by procedure type and no-show patterns.

Agritourism business, guest numbers doubled. Work on positioning and channels, with data analysis used to understand where profitable enquiries actually came from.

The common thread across all four is that none of these results came from a technology project. They came from organizational projects that used AI as a tool. Companies that invert the order, starting from the platform, spend without collecting.

If you recognize your organization in at least two of those patterns, the right next step is not buying software. It is putting someone on measuring, for two weeks, where the time of your regulatory, quality and service teams actually goes, and which two processes account for most of the recoverable hours. That is the work I start with when a company requests a consultation, and it ends with three initiatives ranked by expected return rather than with a platform to purchase. If you want to know which of the three applies to your organization, requesting a consultation is the step that removes months from everything after it.

Self-assessment: is your organization ready

Answer yes or no. One point per yes.

Data and systems

  1. Quality and regulatory records are digital and searchable, not only scanned.
  2. Our document management system exposes an API or supports scheduled export.
  3. There is one authoritative version of core product data, not three spreadsheets.
  4. I can state how many hours a technical file update currently takes.
  5. Complaint records are structured enough to be queried by product family.

Process and organization

  1. There is a recurring meeting where quality and post-market indicators drive decisions.
  2. When an indicator moves, someone owns finding out why.
  3. The process I want to support with AI is documented, at least minimally.
  4. Someone internal can dedicate real time to the project, not leftover time.

Compliance and capability

  1. I know which data cannot leave our perimeter, and who decides that.
  2. Our quality function is involved in technology decisions before procurement, not after.
  3. At least one person internally can challenge a technical vendor rather than defer to one.

Reading the score

Zero to four: you are not ready for projects on quality-relevant data. Start with drafting support and tender content, which need no integration, and use those months to get record access in order.

Five to eight: this is the typical condition of a well-run mid-sized manufacturer. Start with regulatory drafting, literature review and audit preparation, and use those projects to build the foundation for the rest.

Nine to twelve: you have the prerequisites for complaint triage and quality archive retrieval. Your risk is not data quality, it is spreading the investment across too many fronts at once.

The 30, 60, 90 day roadmap

This is the sequence I recommend to a company starting from zero. It is deliberately conservative: the goal of the first ninety days is not to transform the company, it is to produce one measurable result that convinces people it is worth continuing.

Days 1 to 30: measure where the time goes

No purchasing in this phase. Three activities.

The first is time measurement: for two weeks, on a sample of people and with their consent, measure how much time goes into document production, information retrieval, rewriting and administrative work. You need measured data, not a leadership estimate.

The second is a data inventory: where records live, in what format, who can access them, what is exportable and what is locked inside closed systems.

The third is choosing the pilot. One only. The criteria: it must be measurable, it must run on material that already exists, it must not touch a reportability or clinical decision, and it must not span more than one function.

Days 31 to 60: pilot in one function

Implement the chosen case, almost always regulatory drafting support or audit preparation retrieval.

Three rules separate pilots that finish from pilots that drift. Define the number you intend to move before starting. Keep the old process running in parallel for the whole duration. Set a verdict date. A pilot without a verdict date becomes a permanent state, and that is the most common way these projects die without anyone declaring the death.

You also need an internal owner, not a vendor. If nobody inside owns the project, the project does not exist. In most cases the right person is whoever knows the process, not whoever knows the technology.

If the pilot touches a quality-relevant process, the validation documentation starts here, on day 31, not after the pilot succeeds.

Days 61 to 90: measure, then extend or close

Compare the result against the number defined on day 31. If the result is there, extend to other functions, rewrite the procedure and train the people who were not involved. If it is not, close it and write down why: one documented failed pilot is worth more than three open ones.

Only at this point do you choose the second initiative, and it is almost always the one that reuses the data access the first one opened up.

This is also the moment where outside help is worth the money. Not for the technology, which in 2026 is the easy part, but for sequencing: the order in which you take on processes determines whether the second project costs half the first or twice as much. That conversation is worth having with someone who has watched the sequence go wrong elsewhere, before it goes wrong in your building.

The five mistakes I see most often

Starting with the device instead of the back office. The product conversation is the one everybody wants to have and the one with the worst risk-to-return ratio as a first move. The first project exists to build internal credibility, and it builds it on repetitive internal work.

Buying the platform before having the problem. The correct order is problem, data, model, tool. Inverted, it produces active licenses and no usage. It is the most expensive mistake because it is also the hardest to admit after signing.

Treating validation as a later step. Discovering at audit that a quality-relevant process runs on an unvalidated tool means either retrofitting documentation under pressure or removing a system people now depend on. Both are avoidable by deciding on day one.

Calling license distribution adoption. The number of active accounts measures nothing. The indicator is hours recovered on an identified activity, measured before and after.

Ignoring the 2028 design trap. The Annex I extension moved the compliance deadline, not the design deadline. Products reaching the market in 2028 are being designed now, and data provenance decisions taken today determine whether conformity is documentation or redesign.

How this differs by company size

A guide that treats all companies identically is useless. The return on the same initiatives changes sharply with scale.

Under 50 employees. Value is almost entirely in regulatory drafting, tender responses and literature review. Light interventions, no infrastructure, results in weeks. The main risk is buying platforms sized for large groups. The selection criterion is speed of activation, not feature completeness. The reasoning for this scale is in the guide on AI for small business.

Between 50 and 500 employees. This is the band with the highest return relative to investment. Complexity is already high enough that nothing can be run from memory, but the organization is still lean enough to change a process in weeks. Regulatory documentation, audit preparation and complaint triage are the three main levers here.

Above 500 employees. The problem shifts from missing data to data fragmented across systems, sites and functions with divergent practices. Value concentrates in unification and governance, and projects take longer because more functions are involved. The characteristic risk is launching a multi-year program that produces its first useful result after eighteen months, when internal confidence has already evaporated. The countermeasure is the same: one pilot, one function, one verdict date.

The general criterion, valid at every scale: the more repetitive and textual the work, the more automated generation belongs in it. The more it involves judgment and regulatory responsibility, the more the technology should prepare material rather than decide. The overall framing is in the guide on AI implementation for business.

People: the variable that decides more than half the outcome

There is an aspect no vendor deck addresses that weighs more than the technology.

Introducing a system that drafts, summarizes and produces documents touches the professional identity of the people who do that work. In regulatory and quality functions this is sharper than elsewhere, because those professionals are accountable for what they sign, and accountability without control is the definition of an unacceptable position.

The correct message is different, and it has to be true to work: the system exists to remove transcription, retrieval and formatting, which is the part of the job nobody claims as their expertise. The proof comes from the first measured result, not from the initial promise.

Three things that work. Involve two or three people in the selection and configuration, with real authority to say something is unacceptable. Publish the time-recovered number, not the software usage number. Treat objections as calibration material: within two months they become the best source for correcting the system.

There is also a substantive point worth stating plainly. The fear that AI flattens the quality of regulatory work is well founded when it replaces thinking, and unfounded when it removes transcription. That distinction has to be declared out loud, because nobody assumes it and silence gets interpreted in the worst way.

Choosing a vendor without getting hurt

Five questions for anyone proposing an AI project to a device company.

  1. Which cost line or indicator moves, and by how much? If the answer is "efficiency" with no number, the conversation ended there.
  2. Which of my data does it run on, and do I already have it in a usable format? If the project needs structured records you do not have in that form, there is a phase zero nobody quoted, and it often costs as much as the project.
  3. Where is the data processed, who can access it, and how long is it retained? Contractual question, written answer, before signature.
  4. How does this fit my quality system? A vendor who has never heard the question has not worked in this sector, and you will be writing the validation documentation yourself.
  5. What happens when the system produces an error, and who is accountable? There must be a procedure and a defined responsibility, not a commercial reassurance.

One warning sign specific to this market phase: the vendor who demos on their own data and never asks to see yours. The demo always works. The project works only if it survived your real complaint records, with their wrong formats and their exceptions.

What changes in the next twenty-four months

Three movements are already visible, and they are worth watching without waiting for them.

The technical barrier keeps falling, the organizational one does not. Models improve and cost less, but the difficulty of changing how a regulatory department works is unchanged. Competitive advantage shifts toward companies that manage change well, not toward those that adopt earliest.

Regulatory expectations converge. The joint guidance between device and AI regulators in Europe is the first move in a longer sequence. The direction is integration of AI requirements into existing quality processes rather than a parallel compliance track, which is good news for manufacturers who invest in documentation discipline now.

Systems that execute, not only suggest, enter the process. According to Deloitte's analysis of AI agents, adoption is running ahead of governance: only 21% of surveyed organizations say they have a mature governance model in place for agentic AI. The Stanford HAI AI Index 2026 economy chapter confirms it from the other side: agent deployment remains in single digits across nearly every business function. The correct reading is not that it is too early, it is that anyone entering now should enter on internal, reversible processes rather than on anything touching patients, claims or regulatory submissions.

From generated document to approved submission: the last mile

There is a step almost every guide skips, and it is where projects create or destroy value: the moment a model's output becomes an action.

Almost every company I have seen fail on this had perfectly adequate outputs. What was missing was the mechanism that turns a produced document into a decision taken. A summary nobody reads produces nothing. A flagged complaint that changes no activity is an exercise in style.

Building that mechanism is less technological than it sounds and reduces to three elements. Every relevant output has a named recipient, not a distribution list. There is a threshold beyond which action is mandatory rather than discretionary. And there is a fixed weekly moment where someone checks what was done with what was produced the week before.

It sounds trivial, and the hard part is not understanding it. It is sustaining it for six months, when the initial enthusiasm is gone and meetings get shorter. Companies that hold that rhythm get a better return than companies that invested three times as much in technology, and I have seen enough cases to treat it as a rule rather than an observation.

If you are working out where to start, and you have already understood that the problem is not choosing the tool but choosing the sequence, the sensible move is an external measurement of where your regulatory, quality and service hours actually go, before any investment. Two weeks of analysis costs a fraction of a wrong project, and in most cases it completely reorders the priorities leadership had in mind. If you want to set that up for your organization, it is exactly the kind of work I begin every engagement with, and requesting a consultation is the fastest way to start from measurement rather than from procurement.

FAQ

What does AI for medical devices actually mean in practice?

It means two different things that get confused constantly. AI in the product is a model contributing to the device's intended medical purpose, such as detection, measurement or triage, which is a regulated design decision affecting classification, clinical evidence and technical documentation. AI in the company is everything else: regulatory drafting, literature review, complaint triage, audit preparation, tender content and field service support. The near-term measurable return is almost entirely in the second category, and companies that start with the first typically spend eighteen months producing a roadmap instead of a result.

Does the EU AI Act apply to medical devices, and when?

Yes. An AI-enabled medical device falls under both the medical device framework and the AI Act at the same time. Obligations for AI embedded in already-regulated products under Annex I, which covers medical devices, apply from 2 August 2028 after the Digital Omnibus package confirmed by the Council on 29 June 2026 extended the original deadline. Transparency obligations were not postponed and apply now. The Medical Device Coordination Group and the European AI Board published joint guidance in June 2025 clarifying that the two sets of requirements integrate into existing quality processes rather than duplicating them.

What is a Predetermined Change Control Plan and why does it matter?

A PCCP is a plan describing intended future modifications to an AI-enabled device, the methodology to develop and validate them, and an assessment of their impact, submitted and reviewed as part of the original marketing submission. Modifications falling within the approved plan can be implemented without a new submission. The FDA published final guidance on 3 December 2024, and the final version applies to all AI-enabled device software functions rather than only continuously learning systems. Practically, it is the difference between a product you can keep improving and one frozen at its authorized state.

How much does it cost a device company to start with AI?

For a manufacturer between 50 and 500 employees, a serious first project costs between 15,000 and 120,000 USD in the first year depending on the use case. Regulatory drafting support and tender content sit at the low end, complaint triage and quality archive retrieval at the high end. Putting AI inside the device is a different category entirely, starting around 250,000 USD and measured in years rather than months. Model costs are now marginal, since inference cost fell more than 280 times in two years. What weighs is integration and the time of the people changing process.

Can AI be used in complaint handling and post-market surveillance?

Yes, with a specific boundary. AI works well for reading unstructured complaint text, extracting structured fields, proposing a failure classification and flagging items that look reportable. It should not make the reportability decision itself. Two safeguards are mandatory: a confidence threshold below which items route to a human queue, and a fixed control sample re-checked manually every week, because these systems degrade quietly when incoming formats change. If the process is quality-relevant, the system also needs to be handled as a validated system from day one.

Do AI tools used in a regulated company need validation?

If the tool participates in a process governed by your quality system, yes. That means documented intended use, validation evidence, change control and a defined behavior when the system fails or is unavailable. The common failure mode is a team quietly adopting a general-purpose assistant for drafting or triage, generating real value for six months, then discovering at audit that a quality-relevant process ran on an undocumented tool. Deciding on day one whether a use is quality-relevant, and writing down the reasoning either way, costs far less than retrofitting that documentation under audit pressure.

Where should a medical device company start if nobody has used AI before?

With regulatory document drafting and audit preparation retrieval. Both run on material that already exists, need no integration with production systems, touch no clinical or reportability decision, and produce a visible result within weeks. Before that, spend two weeks measuring where regulatory, quality and service hours actually go, because in most companies those four numbers do not exist and without them nobody can say whether an intervention worked. Starting with AI inside the device, or with anything touching patient-facing decisions, means taking on integration, evidence and regulatory risk simultaneously on the first attempt.