ISSN (Online): 2321-3418
server-injected
Engineering and Computer Science
Open Access

AI-Assisted Quality and Regulatory Compliance Framework for Robotic Systems Used in Minimally Invasive Surgery

DOI: 10.18535/ijsrm/v13i08.ec07· Pages: 2615-2624· Vol. 14, No. 07, (2026)· Published: August 28, 2025
PDFAuto
Views: 7 PDF downloads: 6

Abstract

Robotic systems used in minimally invasive surgery are becoming increasingly dependent on artificial intelligence for surgical planning, image interpretation, workflow recognition, technical skill assessment, decision support, and partial automation. Despite these advances, conventional medical-device quality systems do not fully address the changing behaviour, data dependence, explainability limitations, cybersecurity exposure, and post-deployment performance variation associated with artificial intelligence-enabled surgical robots. This paper proposes an AI-Assisted Quality and Regulatory Compliance Framework for the development, validation, deployment, and continuous monitoring of robotic systems used in minimally invasive surgery. The framework integrates six interrelated domains: governance and accountability, risk-based design control, data and algorithm quality, clinical and technical validation, regulatory evidence management, and post-market surveillance. It positions artificial intelligence not only as a component requiring control but also as a quality-assurance instrument capable of identifying anomalies, monitoring surgical performance, detecting deviations, supporting traceability, and generating compliance evidence. The proposed framework adopts a lifecycle approach in which quality, safety, ethical responsibility, human oversight, and regulatory requirements are considered from initial design through long-term clinical use. It also introduces control requirements aligned with different levels of robotic autonomy and emphasizes the importance of representative data, surgeon supervision, change control, transparent performance reporting, and continuous algorithm surveillance. The framework offers manufacturers, healthcare institutions, regulators, surgeons, and quality professionals a structured basis for managing the risks of increasingly intelligent surgical robots while supporting responsible innovation. Its implementation could improve patient safety, strengthen regulatory readiness, enhance clinical confidence, and reduce the gap between rapid technological development and existing medical-device quality practices.

Keywords

artificial intelligence robotic surgery minimally invasive surgery quality management regulatory compliance surgical robotics medical devices algorithm validation post-market surveillance human oversight surgical data science autonomous systems

1. Introduction

Robotic systems have significantly influenced minimally invasive surgery by improving instrument dexterity, motion scaling, visualization, access to difficult anatomical regions, and the precision of surgical movements. The combination of robotics and artificial intelligence has expanded these capabilities beyond mechanical assistance toward image-guided navigation, recognition of operative phases, automated performance assessment, intraoperative decision support, and task-level autonomy. Andras et al. (2020) observed that the convergence of artificial intelligence and robotics is changing the operating room by transforming how surgical information is interpreted and how procedures are performed.

Artificial intelligence-assisted surgery includes a broad range of functions. These include the analysis of endoscopic video, identification of anatomical structures, prediction of procedural risks, recognition of surgical instruments, assessment of surgeon activity, and support for robotic control. Bodenstedt et al. (2020) argued that artificial intelligence could improve surgical decision-making and procedural standardization, although its clinical value depends on data quality, validation, workflow integration, and acceptance by surgeons. Machine learning can also support the optimization of robotic procedures by examining surgical motion, identifying technical patterns, and predicting performance outcomes (Ma et al., 2020).

The development of intelligent robotic surgery has been supported by research platforms such as the da Vinci Research Kit, which has enabled researchers to investigate computer vision, robotic control, automation, human-machine interaction, and surgical data analysis. D’Ettorre et al. (2021) showed that research using this platform has contributed to progress in areas such as autonomous manipulation, image guidance, force estimation, and skills assessment. However, the transition from controlled laboratory studies to routine clinical use introduces substantial safety, quality, and regulatory concerns.

Artificial intelligence systems are influenced by the data on which they are trained, the conditions under which they are deployed, and the changes made to their software after release. A model may perform well during development but produce unreliable outputs when exposed to new patient populations, surgical techniques, instruments, lighting conditions, or clinical environments. Hashimoto et al. (2018) described artificial intelligence in surgery as offering considerable promise while also introducing risks involving bias, weak interpretability, data privacy, inappropriate reliance, and uncertain responsibility.

These problems are especially important in minimally invasive surgery because surgical robots operate in complex, time-sensitive environments in which small technical failures may affect patient safety. The expansion of robotic autonomy adds further difficulty. A systematic review by Lee et al. (2024) found that most identified FDA-cleared surgical robots remained at a low level of autonomy, while only a small proportion demonstrated conditional autonomy. The review identified 49 surgical robots, of which most were classified at Level 1, indicating that present systems still depend heavily on human control. Nevertheless, experimental research has demonstrated that autonomous robots can perform selected soft-tissue surgical tasks. Saeidi et al. (2022), for example, demonstrated autonomous robotic laparoscopic intestinal anastomosis in experimental settings.

The increasing intelligence and autonomy of robotic systems create a gap between technological capability and traditional quality-management practices. Conventional medical-device quality systems emphasize design control, verification, validation, corrective action, supplier control, complaint handling, and risk management. These remain essential, but they may not sufficiently address algorithmic bias, data drift, model retraining, explainability, changing performance, or human-machine interaction.

This paper proposes an AI-Assisted Quality and Regulatory Compliance Framework for robotic systems used in minimally invasive surgery. The framework is designed to support manufacturers, healthcare institutions, regulators, quality professionals, software developers, and surgeons. Its primary objectives are to:

  • integrate artificial intelligence governance into the medical-device quality lifecycle;

  • connect algorithm development with clinical, technical, ethical, and regulatory requirements;

  • establish controls proportionate to the autonomy and risk of the robotic function;

  • use artificial intelligence to strengthen quality monitoring and regulatory evidence management; and

  • support continuous surveillance throughout the operational life of the robotic system.

2. Quality and Regulatory Challenges in AI-Enabled Surgical Robotics

2.1 Complexity of AI-enabled robotic systems

An AI-enabled surgical robot is not a single technology. It is a combination of hardware, software, sensors, imaging devices, control systems, user interfaces, clinical procedures, datasets, and human decisions. A failure may therefore arise from several sources, including inaccurate sensor measurements, software errors, communication delays, mechanical faults, poor image quality, biased training data, incorrect model outputs, or inappropriate human responses.

The increasing use of surgical data science adds another layer of complexity. Surgical data science involves the systematic collection, organization, analysis, and interpretation of information generated before, during, and after surgery. Maier-Hein et al. (2022) explained that successful clinical translation requires more than accurate algorithms. It also requires standardized data, interdisciplinary collaboration, clinical relevance, workflow compatibility, ethical governance, and reproducible evaluation.

The quality system must therefore control both the physical robot and the information ecosystem surrounding it. Data pipelines, annotation procedures, software updates, model parameters, clinical interfaces, and user behaviour can all influence system performance. Quality assurance cannot be limited to testing the robot at the end of development. It must be embedded throughout the entire product lifecycle.

2.2 Data quality and representativeness

The reliability of artificial intelligence depends heavily on the quality of the data used for training, validation, and testing. Surgical datasets may vary according to hospital, surgeon experience, patient anatomy, procedure type, device generation, imaging equipment, recording quality, and institutional practice. A model trained on data from a limited number of expert surgeons or high-resource hospitals may not perform equally well in other environments.

Video-based surgical models illustrate this challenge. Kiyasseh et al. (2023) developed a vision transformer capable of decoding surgeon activity from surgical videos. Such systems can support activity recognition, workflow monitoring, and performance assessment. However, their effectiveness depends on accurate labelling, representative examples, consistent video quality, and validation across different surgeons and clinical sites.

Data quality controls should address completeness, accuracy, relevance, traceability, class balance, annotation consistency, privacy protection, and population representation. The source of each dataset, its collection conditions, inclusion criteria, exclusions, preprocessing operations, and known limitations should be documented. Data used for testing should remain sufficiently independent from data used for model development.

2.3 Validation and clinical evidence

Technical accuracy alone does not demonstrate that an AI-enabled surgical robot is safe or clinically beneficial. A model may correctly recognize an operative phase but fail to provide information at the right time. An autonomous function may perform well in a controlled laboratory but become unreliable when anatomical movement, bleeding, smoke, instrument occlusion, or unexpected tissue variation occurs.

Validation should therefore occur at several levels:

  • software and algorithm verification;

  • subsystem and hardware integration testing;

  • simulated-use evaluation;

  • usability and human-factors testing;

  • preclinical testing;

  • clinical performance evaluation; and

  • long-term post-market monitoring.

Marcus et al. (2024) emphasized that surgical robots require structured evaluation across development, comparative assessment, and long-term clinical use. The IDEAL framework recognizes that robotic technologies evolve over time and that their evaluation must consider learning curves, clinical context, operator experience, system modifications, and broader healthcare effects.

Training and simulation also form part of validation and safe implementation. Moglia et al. (2016) found that virtual-reality simulators can support training in robot-assisted surgery, although evidence quality and validation methods vary. Simulator-based training can help confirm that surgeons understand system limitations, emergency procedures, interface behaviour, and appropriate responses to AI recommendations.

2.4 Ethical, legal, and accountability concerns

AI-assisted surgery raises questions concerning responsibility when an algorithmic recommendation or robotic action contributes to an adverse outcome. Potentially responsible parties may include the surgeon, hospital, manufacturer, software developer, data provider, or maintenance organization. O’Sullivan et al. (2019) identified legal and regulatory concerns involving liability, privacy, medical-device law, negligence, transparency, and accountability.

Morris et al. (2023) similarly emphasized that the use of artificial intelligence in surgery involves ethical, legal, and financial implications. These include informed consent, ownership of surgical data, unequal access, explainability, reimbursement, responsibility for errors, and the cost of implementing intelligent technologies.

Accountability becomes more difficult as autonomy increases. In surgeon-controlled systems, the surgeon directly performs the procedure through the robot. In partially autonomous systems, the machine may complete selected tasks under supervision. In conditionally autonomous systems, the system may act independently within defined boundaries. Each level requires clear identification of who authorizes the action, who monitors performance, when human intervention is required, and how decisions are recorded.

2.5 Post-deployment change and performance deterioration

Unlike fixed mechanical devices, AI models may be retrained, recalibrated, or updated after deployment. Even when the model remains unchanged, its performance may deteriorate because the clinical environment changes. New instruments, different cameras, revised operating techniques, new patient groups, or altered data-processing procedures can produce data drift.

Post-market monitoring must therefore evaluate algorithm performance rather than only mechanical reliability and complaint frequency. Relevant indicators include false warnings, missed detections, override rates, unexpected disengagement, task-completion failure, variation across patient groups, and differences across clinical sites.

Table 1 Principal quality and regulatory risks in AI-enabled surgical robotics
Risk area Example of potential failure Possible consequence Required quality response
Training-data quality Incomplete, inaccurately labelled, or unrepresentative surgical data Biased or unreliable model output Data qualification, annotation review, dataset version control, representation analysis
Algorithm performance Reduced accuracy under smoke, bleeding, occlusion, or unusual anatomy Incorrect guidance or delayed intervention Stress testing, subgroup validation, uncertainty thresholds, fail-safe design
Hardware-software integration Delay or error between AI output and robotic control Unintended robotic movement or workflow interruption Integration verification, latency testing, fault injection, interface control
Human factors Surgeon misunderstands AI recommendation or system status Overreliance, delayed correction, or misuse Usability testing, clear interface design, competency assessment, training
Cybersecurity Unauthorized access or manipulation of software and data Loss of confidentiality, integrity, or system availability Access control, encryption, vulnerability monitoring, incident response
Model update Uncontrolled software or algorithm modification Changed clinical performance after release Formal change control, revalidation, impact assessment, regulatory review
Accountability Unclear responsibility for AI-assisted decisions Delayed response, legal dispute, weak corrective action Defined responsibility matrix, audit trails, escalation procedures
Post-market drift Performance changes across sites or over time Increasing rate of errors or missed events Continuous monitoring, site comparison, drift detection, corrective action

3. Proposed AI-Assisted Quality and Regulatory Compliance Framework

The proposed framework consists of six integrated domains. These domains should not be treated as independent activities. Each domain produces evidence that supports the others and contributes to an auditable lifecycle record.

3.1 Governance, accountability, and human oversight

The first domain establishes responsibility for the development, approval, deployment, monitoring, and modification of the AI-enabled robotic system. A multidisciplinary governance body should include representatives from clinical surgery, quality assurance, regulatory affairs, software engineering, robotics, cybersecurity, data science, human factors, ethics, and patient safety.

The governance structure should define:

  • the intended clinical purpose of each AI function;

  • the approved level of autonomy;

  • the responsibilities of the surgeon and robotic system;

  • criteria for human intervention;

  • authority to approve datasets, models, and updates;

  • processes for investigating failures;

  • responsibility for regulatory communication; and

  • rules for the retention and use of surgical data.

Human oversight must be designed according to the risk of the function. An AI system that labels instruments in recorded videos presents a different risk from one that controls tissue manipulation in real time. Han et al. (2022) described the progression of robotic surgery from supervised systems toward increasingly autonomous approaches. This progression requires stronger control mechanisms as robotic authority expands.

For high-risk functions, the system should provide clear status information, confidence estimates, warnings, override controls, and safe transition to manual operation. Human oversight must be meaningful rather than symbolic. The surgeon must have sufficient information, time, training, and physical ability to intervene.

3.2 Risk-based design and development control

The second domain integrates AI-specific risks into design control. At the beginning of development, the manufacturer should define the intended use, target procedure, users, patient population, operating environment, contraindications, expected benefits, and reasonably foreseeable misuse.

Risk analysis should examine the full chain between input data and clinical action. This includes:

  1. data acquisition;

  2. sensor or image processing;

  3. algorithmic interpretation;

  4. communication of the output;

  5. robotic or human response; and

  6. effect on the patient.

The safety assessment should consider both component failure and interaction failure. A technically correct prediction may still create harm if it is presented late, displayed ambiguously, or applied outside its intended clinical context. Bodenstedt et al. (2020) emphasized that the practical value of AI-assisted surgery depends on its integration into surgical workflow and its ability to operate reliably under clinical conditions.

Design controls should include documented user needs, system requirements, software requirements, architecture specifications, interface requirements, verification protocols, validation plans, and traceability records. Requirements should be measurable. For example, a requirement should specify the minimum acceptable sensitivity, maximum response time, permitted uncertainty, and expected behaviour when input quality becomes inadequate.

3.3 Data and algorithm quality management

The third domain controls data and model development. Every dataset should have a documented purpose and should be assessed before use. The assessment should cover provenance, patient population, surgical procedure, data format, collection device, surgeon characteristics, consent status, missing information, annotation method, and known limitations.

A formal data lifecycle should include:

  • data acquisition and authorization;

  • de-identification and privacy protection;

  • quality screening;

  • annotation and adjudication;

  • preprocessing;

  • dataset partitioning;

  • model training;

  • independent validation;

  • secure storage;

  • version control; and

  • Controlled retirement.

Algorithm development should be reproducible. Model architecture, software libraries, parameter settings, training procedures, performance metrics, and computational environments should be documented. This is particularly important for complex models used to interpret surgical videos. Kiyasseh et al. (2023) demonstrated that transformer-based models can identify surgeon activities from video, but reliable use requires careful evaluation across procedures, operators, and environments.

Performance should be evaluated with metrics relevant to the intended task. Accuracy alone may be misleading when outcome classes are imbalanced. Depending on the application, evaluation may include sensitivity, specificity, precision, recall, calibration, latency, task success, false-alert rate, override frequency, and subgroup performance.

Bias analysis should examine whether performance differs according to patient characteristics, anatomical variation, hospital, device configuration, or surgeon experience. When meaningful differences are detected, they should be investigated before deployment.

3.4 Verification, validation, and clinical evaluation

The fourth domain establishes evidence that the system meets its requirements and provides clinical value. Verification confirms that the system has been built according to its specifications. Validation confirms that it performs safely and effectively for its intended users and use conditions.

Validation should progress from lower-risk environments to realistic clinical conditions. Early stages may involve software testing, benchtop testing, recorded video, synthetic data, phantom models, virtual reality, or cadaveric studies. Later stages may include animal studies, controlled clinical investigations, comparative evaluations, and post-market studies.

The evaluation strategy should be proportionate to the autonomy level. A decision-support model may require evidence that its recommendations are accurate, timely, interpretable, and appropriately used by surgeons. An autonomous task requires additional evidence regarding motion safety, tissue interaction, emergency stopping, recovery from failure, boundary recognition, and human takeover.

The work of Saeidi et al. (2022) illustrates both the potential and the evidentiary demands of autonomous surgery. The reported robotic system performed laparoscopic small-bowel anastomosis in phantom and in vivo tissues, demonstrating progress in autonomous soft-tissue surgery. However, movement from an experimental demonstration to general clinical use would require broader validation across anatomies, operating conditions, teams, institutions, and unexpected events.

Technical skill assessment is another important application of AI. Pedrett et al. (2023) reported that artificial intelligence can support assessment of technical performance in minimally invasive surgery. Such systems may improve training and quality monitoring, but they require clear definitions of competence, validated outcome measures, transparent scoring, and protection against unfair assessment.

Table 2 Validation requirements according to robotic autonomy
Autonomy category Typical system role Minimum validation emphasis Required human control
Assistance Improves visualization, filters motion, or supports instrument positioning Mechanical accuracy, image quality, usability, latency, reliability Continuous direct surgeon control
Decision support Detects anatomy, predicts risk, or recommends an action Diagnostic performance, calibration, interpretability, alert timing, subgroup testing Surgeon reviews and accepts or rejects output
Task automation Completes a defined surgical subtask Task success, boundary detection, tissue safety, interruption handling, recovery testing Active supervision with immediate override
Conditional autonomy Performs a task independently under defined conditions Environmental limits, uncertainty detection, fail-safe behaviour, takeover testing, extensive clinical evidence Surgeon available to intervene when requested
Higher autonomy Plans or performs substantial procedural activity System-level safety, broad generalization, ethical acceptability, continuous monitoring, exceptional-event handling Clearly defined supervisory and emergency authority

3.5 Regulatory evidence and traceability management

The fifth domain connects quality activities with regulatory documentation. Every important claim about the device should be supported by traceable evidence. The intended use should be linked to clinical requirements. Requirements should be linked to identified risks, design outputs, verification tests, validation results, and post-market indicators.

An AI-assisted regulatory evidence system can organize and examine these relationships. Natural-language processing and rule-based tools can identify missing documents, inconsistent terminology, expired approvals, incomplete risk controls, or requirements without corresponding test evidence. However, the final regulatory decision should remain under qualified human authority.

The regulatory evidence package should include:

  • intended-use and autonomy statements;

  • system and software descriptions;

  • dataset documentation;

  • algorithm-development records;

  • risk-management documentation;

  • cybersecurity evidence;

  • usability and human-factors reports;

  • verification and validation results;

  • clinical evaluation;

  • training requirements;

  • change-management plans; and

  • post-market surveillance procedures.

The IDEAL framework proposed by Marcus et al. (2024) is particularly relevant because it treats surgical robotics as an evolving clinical innovation requiring evidence across development, evaluation, and long-term monitoring. The quality system should therefore preserve evidence from early development rather than assembling regulatory documentation only at the point of submission.

Regulatory claims must also accurately reflect actual system capability. Lee et al. (2024) reported differences between machine-learning capabilities formally recognized through regulatory records and capabilities presented in some marketing materials. This finding supports the need for consistent descriptions across regulatory submissions, technical documentation, user manuals, promotional materials, and clinical training.

3.6 Post-market surveillance and continuous improvement

The sixth domain establishes continuous monitoring after clinical deployment. Conventional complaint systems are often reactive because they depend on users identifying and reporting a problem. AI-assisted surveillance can provide earlier indications by analysing performance logs, surgical video, system warnings, override events, technical errors, and outcome patterns.

Post-market indicators may include:

  • frequency of AI recommendations;

  • acceptance and rejection rates;

  • false-positive and false-negative events;

  • surgeon override frequency;

  • unplanned system disengagement;

  • emergency conversion to manual surgery;

  • procedure completion rates;

  • technical-error reports;

  • performance variation across hospitals;

  • subgroup performance;

  • cybersecurity events; and

  • patient outcomes associated with system use.

The surveillance system should define thresholds for review, corrective action, model suspension, retraining, or regulatory notification. Automated alerts should not directly determine that a system is safe. They should identify signals for investigation by qualified personnel.

When an update is proposed, the organization should assess whether it affects intended use, clinical performance, risk controls, cybersecurity, interoperability, user training, or regulatory status. The updated model should not be released until the required verification and validation activities are completed.

Table 3 Core elements of the proposed framework
Autonomy category Typical system role Minimum validation emphasis Required human control
Assistance Improves visualization, filters motion, or supports instrument positioning Mechanical accuracy, image quality, usability, latency, reliability Continuous direct surgeon control
Decision support Detects anatomy, predicts risk, or recommends an action Diagnostic performance, calibration, interpretability, alert timing, subgroup testing Surgeon reviews and accepts or rejects output
Task automation Completes a defined surgical subtask Task success, boundary detection, tissue safety, interruption handling, recovery testing Active supervision with immediate override
Conditional autonomy Performs a task independently under defined conditions Environmental limits, uncertainty detection, fail-safe behaviour, takeover testing, extensive clinical evidence Surgeon available to intervene when requested
Higher autonomy Plans or performs substantial procedural activity System-level safety, broad generalization, ethical acceptability, continuous monitoring, exceptional-event handling Clearly defined supervisory and emergency authority
Autonomy category Typical system role Minimum validation emphasis Required human control

4. Framework Implementation and Operational Control

4.1 Lifecycle implementation process

Implementation should begin with a gap assessment comparing the proposed framework with the organization’s existing quality-management system. The organization should identify which AI-specific controls are already present and which require development.

The implementation process can follow six stages.

Stage 1: Define the system and intended use.

The manufacturer should clearly describe the surgical procedure, clinical objective, users, patient population, operating environment, data inputs, robotic outputs, and level of autonomy.

Stage 2: Establish multidisciplinary governance.

A governance committee should approve the intended use, autonomy level, dataset strategy, validation plan, risk-acceptance criteria, and update process.

Stage 3: Build traceable data and development controls.

Dataset versions, annotations, model versions, requirements, software changes, and test results should be stored in a controlled system.

Stage 4: Conduct risk-based verification and validation.

Testing should examine normal operation, foreseeable misuse, degraded input conditions, rare events, hardware-software interactions, cybersecurity threats, and human takeover.

Stage 5: Prepare regulatory and deployment evidence.

The organization should confirm that claims, labels, training materials, risk controls, and clinical evidence are consistent.

Stage 6: Activate continuous monitoring.

Post-market indicators, alert thresholds, review responsibilities, reporting schedules, and corrective-action procedures should be defined before commercial deployment.

4.2 Human factors and surgeon competency

A surgical robot can be technically reliable but unsafe if its interface is confusing or if users misunderstand its limitations. Human-factors engineering should therefore examine how surgeons interpret AI outputs, warnings, confidence information, system states, and requests for intervention.

Training should cover normal operation, system limitations, likely failure modes, emergency procedures, manual takeover, cybersecurity responsibilities, and appropriate reliance on AI. Virtual-reality simulation can support structured training without exposing patients to unnecessary risk. Moglia et al. (2016) found that virtual-reality simulators have important potential for robot-assisted surgical education, although validation and transfer of training to clinical performance must be demonstrated.

Competency should be assessed rather than assumed after attendance at training. AI-based technical skill assessment may help identify performance patterns and training needs, as discussed by Pedrett et al. (2023). However, such assessments should not be used as the sole basis for high-stakes decisions unless their validity, fairness, and reliability have been established.

4.3 Quality indicators and acceptance thresholds

Each AI function should have predefined quality indicators. These indicators should reflect technical performance, clinical performance, user interaction, and patient safety.

For an anatomical recognition model, indicators might include sensitivity, precision, time to detection, failure under occlusion, and performance across patient subgroups. For an autonomous suturing task, indicators might include placement accuracy, tissue damage, completion time, interruption rate, recovery success, and frequency of human takeover.

Acceptance thresholds should be established before validation begins. Changing the success criteria after results are known can introduce bias. Thresholds should be clinically justified rather than selected only because they are technically achievable.

4.4 Audit and corrective-action processes

Internal audits should assess whether AI-specific procedures are followed and whether the evidence remains complete. Audit activities should examine dataset approval, annotation quality, access control, model versioning, test reproducibility, validation independence, deployment records, update authorization, and post-market review.

Corrective and preventive action should be activated when performance falls below an approved threshold, an unexpected risk emerges, or a regulatory requirement is not met. The investigation should determine whether the problem arose from data, algorithm design, hardware, workflow, training, environment, or human interaction.

Corrective action may involve updating instructions, modifying the interface, retraining users, restricting the intended use, improving data, revising the model, strengthening risk controls, or temporarily suspending the AI function.

5. Discussion

The proposed framework addresses an important limitation in the governance of intelligent surgical systems. AI should not be treated only as an additional software module within a conventional robot. Its performance depends on data, clinical context, model design, user behaviour, and continuous change. These characteristics require a lifecycle quality model.

The framework also recognizes the dual role of artificial intelligence. First, AI is a regulated function that must be validated and controlled. Second, AI can assist the quality system by detecting anomalies, organizing evidence, identifying missing documentation, monitoring workflow, and supporting early risk detection. This dual role can improve efficiency, but it does not remove the need for professional judgement.

A major strength of the framework is its connection between autonomy and evidence. Lee et al. (2024) showed that most currently cleared surgical robots remain primarily assistive, although higher levels of autonomy are emerging. As autonomy increases, the consequences of algorithmic error become more direct. Validation, monitoring, explainability, and takeover requirements should therefore become more demanding.

The framework is also consistent with the progression described by Han et al. (2022), who examined the movement from supervised robotic surgery toward fully autonomous approaches. This progression should not be viewed simply as a technical scale. Each level changes the distribution of responsibility between the surgeon and the system.

Research platforms have accelerated innovation, as demonstrated by D’Ettorre et al. (2021), but research success does not automatically establish clinical readiness. Experimental systems are often tested with selected tasks, controlled data, expert teams, and limited variation. Clinical deployment requires evidence across broader conditions, users, institutions, and patient populations.

Similarly, the autonomous laparoscopic work reported by Saeidi et al. (2022) represents an important technical achievement, but widespread clinical translation would require extensive controls for unexpected anatomy, bleeding, motion, equipment differences, and emergency intervention. The proposed framework provides a structure for organizing such evidence.

Surgical data science can strengthen these efforts by supporting reproducible evaluation, standardized datasets, and performance monitoring. Maier-Hein et al. (2022) argued that clinical translation requires collaboration among surgeons, engineers, computer scientists, institutions, and regulators. The proposed governance model reflects this need.

Ethical and legal concerns remain difficult to resolve. O’Sullivan et al. (2019) and Morris et al. (2023) showed that AI and autonomous surgery create questions about liability, privacy, fairness, transparency, and financial access. A quality framework cannot independently settle legal responsibility, but it can produce clear documentation of system behaviour, decisions, warnings, overrides, updates, and professional responsibilities. Such evidence can improve accountability and support investigation.

The framework has several limitations. It is conceptual and has not yet been evaluated within a specific manufacturer, hospital, or regulatory submission. Its implementation requirements may differ across jurisdictions and surgical specialties. Some performance indicators may also be difficult to standardize because surgical procedures vary substantially across patients and institutions.

Future research should test the framework through case studies involving different robotic systems and autonomy levels. Researchers should examine whether the framework improves defect detection, traceability, regulatory readiness, validation quality, post-market signal identification, and response time. Additional work is also required to establish internationally accepted definitions of surgical autonomy, minimum dataset documentation, algorithm-change categories, and performance-reporting standards.

6. Conclusion

Artificial intelligence is expanding the capabilities of robotic systems used in minimally invasive surgery. It can support image interpretation, activity recognition, surgical planning, technical skill assessment, decision support, and autonomous task performance. However, these capabilities also create risks that are not fully addressed by conventional medical-device quality processes.

The AI-Assisted Quality and Regulatory Compliance Framework proposed in this paper integrates governance, risk-based design control, data and algorithm quality, technical and clinical validation, regulatory evidence management, and post-market surveillance. It applies a lifecycle approach in which AI performance is continuously examined from initial development through long-term clinical use.

The framework emphasizes that regulatory compliance should not be treated as a documentation exercise completed after system development. Safety, evidence generation, traceability, human oversight, cybersecurity, and ethical responsibility should be built into the design and operation of the robotic system.

Successful implementation will require collaboration among surgeons, manufacturers, software engineers, data scientists, regulators, quality professionals, healthcare institutions, and patients. As robotic systems move toward greater autonomy, the strength of their quality and regulatory controls must increase accordingly. A structured framework can support responsible innovation while protecting patient safety, improving clinical confidence, and ensuring that technological progress remains aligned with healthcare obligations.

References

  1. Andras, I., Mazzone, E., van Leeuwen, F. W., De Naeyer, G., van Oosterom, M. N., Beato, S., ... & Mottrie, A. (2020). Artificial intelligence and robotics: a combination that is changing the operating room. World journal of urology, 38(10), 2359-2366. DOI ↗ Google Scholar ↗
  2. Bodenstedt, S., Wagner, M., Müller-Stich, B. P., Weitz, J., & Speidel, S. (2020). Artificial intelligence-assisted surgery: potential and challenges. Visceral Medicine, 36(6), 450-455. Google Scholar ↗
  3. D’Ettorre, C., Mariani, A., Stilli, A., y Baena, F. R., Valdastri, P., Deguet, A., ... & Stoyanov, D. (2021). Accelerating surgical robotics research: A review of 10 years with the da vinci research kit. IEEE Robotics & Automation Magazine, 28(4), 56-78. Google Scholar ↗
  4. Hashimoto, D. A., Rosman, G., Rus, D., & Meireles, O. R. (2018). Artificial intelligence in surgery: promises and perils. Annals of surgery, 268(1), 70-76. DOI ↗ Google Scholar ↗
  5. Han, J., Davids, J., Ashrafian, H., Darzi, A., Elson, D. S., & Sodergren, M. (2022). A systematic review of robotic surgery: From supervised paradigms to fully autonomous robotic approaches. The International Journal of Medical Robotics and Computer Assisted Surgery, 18(2), e2358. Google Scholar ↗
  6. Kiyasseh, D., Ma, R., Haque, T. F., Miles, B. J., Wagner, C., Donoho, D. A., ... & Hung, A. J. (2023). A vision transformer for decoding surgeon activity from surgical videos. Nature biomedical engineering, 7(6), 780-796. DOI ↗ Google Scholar ↗
  7. Lee, A., Baker, T. S., Bederson, J. B., & Rapoport, B. I. (2024). Levels of autonomy in FDA-cleared surgical robots: a systematic review. NPJ Digital Medicine, 7(1), 103. DOI ↗ Google Scholar ↗
  8. Ma, R., Vanstrum, E. B., Lee, R., Chen, J., & Hung, A. J. (2020). Machine learning in the optimization of robotics in the operative field. Current opinion in urology, 30(6), 808-816. DOI ↗ Google Scholar ↗
  9. Maier-Hein, L., Eisenmann, M., Sarikaya, D., März, K., Collins, T., Malpani, A., ... & Speidel, S. (2022). Surgical data science–from concepts toward clinical translation. Medical image analysis, 76, 102306. DOI ↗ Google Scholar ↗
  10. Marcus, H. J., Ramirez, P. T., Khan, D. Z., Layard Horsfall, H., Hanrahan, J. G., Williams, S. C., ... & Additional collaborators Sedrakyan Art 63 Horowitz Joel 64 Paez Arsenio 65. (2024). The IDEAL framework for surgical robotics: development, comparative evaluation and long-term monitoring. Nature medicine, 30(1), 61-75. DOI ↗ Google Scholar ↗
  11. Moglia, A., Ferrari, V., Morelli, L., Ferrari, M., Mosca, F., & Cuschieri, A. (2016). A systematic review of virtual reality simulators for robot-assisted surgery. European urology, 69(6), 1065-1080. DOI ↗ Google Scholar ↗
  12. Morris, M. X., Song, E. Y., Rajesh, A., Asaad, M., & Phillips, B. T. (2023). Ethical, legal, and financial considerations of artificial intelligence in surgery. The American Surgeon, 89(1), 55-60. DOI ↗ Google Scholar ↗
  13. O'Sullivan, S., Nevejans, N., Allen, C., Blyth, A., Leonard, S., Pagallo, U., ... & Ashrafian, H. (2019). Legal, regulatory, and ethical frameworks for development of standards in artificial intelligence (AI) and autonomous robotic surgery. The international journal of medical robotics and computer assisted surgery, 15(1), e1968. DOI ↗ Google Scholar ↗
  14. Pedrett, R., Mascagni, P., Beldi, G., Padoy, N., & Lavanchy, J. L. (2023). Technical skill assessment in minimally invasive surgery using artificial intelligence: a systematic review. Surgical endoscopy, 37(10), 7412-7424. DOI ↗ Google Scholar ↗
  15. Saeidi, H., Opfermann, J. D., Kam, M., Wei, S., Léonard, S., Hsieh, M. H., ... & Krieger, A. (2022). Autonomous robotic laparoscopic surgery for intestinal anastomosis. Science robotics, 7(62), eabj2908. Google Scholar ↗
Author details