Accuracy and Safety of Large Language Models in Medical Diagnostics and Clinical Decision-Making: A Structured Narrative Review
Shubhreet Bhullar (1,2,*), Ahmad Elhaija (2,3)
-
University of California, Los Angeles, 330 De Neve Drive, Los Angeles, CA 90095, USA, thesbhullar22@ucla.edu
-
International Healthcare Organization, Los Angeles, CA, USA
-
David Geffen School of Medicine, University of California, Los Angeles, 10833 Le Conte Ave, Los Angeles, CA 90095, USA
*Corresponding author.
Published: September 2026
DOI: https://doi.org/10.58417/QLAD3062
Abstract
Background: Large language models (LLMs) are moving from experimental demonstrations into clinical decision support, yet much of the literature still emphasizes static examinations rather than the sequential and safety-critical work of patient care. This review evaluates LLM accuracy in differential diagnosis, diagnostic interviewing, triage, treatment planning, guideline adherence, and specialty-specific clinical reasoning.
Methods: A structured narrative synthesis was conducted using a predefined corpus of 20 peer-reviewed clinical sources published through 2026. Evidence was organized by task, model, specialty, comparator, and safety outcome. Methodological quality and risk of bias were assessed using design-appropriate domains adapted from QUADAS-2, ROBINS-I, and AMSTAR 2.
Results: Performance was strongly task dependent. Advanced models often performed well in final-diagnosis selection, medical knowledge, conversational interviewing, and structured multi-agent reasoning. Accuracy was less reliable during early differential formation, diagnostic test selection, clinical calculator use, emergency triage, and management under incomplete or distracting information. Specialty-specific systems showed promise in oncology, radiology, cardiology, and psychiatry, but results varied across subspecialties and study designs. Hallucinated facts, fabricated citations, prompt sensitivity, overconfidence, and adversarial susceptibility remained clinically important hazards. Fourteen of nineteen empirically assessable sources were judged to have moderate overall concern and five high concern; one conceptual source was not formally rated.
Conclusions: Current LLMs can support clinical reasoning, but their performance remains uneven across the diagnostic and management workflow. The evidence favors physician augmentation rather than autonomous use. Safe deployment will require prospective validation, standardized outcomes, transparent model reporting, robust hallucination controls, and workflows that preserve meaningful clinician oversight.
Keywords: large language models; clinical decision support; diagnostic accuracy; medical artificial intelligence; triage; hallucination; human-AI collaboration
Introduction
Artificial intelligence has contributed to medicine for decades through expert systems, risk-prediction tools, image classifiers, and algorithmic decision support. Transformer-based large language models represent a more recent shift. Unlike systems designed for a single narrow task, LLMs can interpret and generate natural language across broad domains. General-purpose models such as ChatGPT, GPT-4, Claude, and Gemini, together with medically adapted systems such as Med-PaLM, can summarize records, synthesize clinical narratives, propose diagnoses, and communicate recommendations in fluent prose. These capabilities have prompted growing interest in their use at the point of care.
Early medical evaluations showed that model scale, instruction tuning, and domain adaptation could improve performance on medical question-answering benchmarks. Med-PaLM encoded substantial clinical knowledge but continued to show deficits in factuality, precision, potential harm, and bias when clinicians evaluated its answers (1). Med-PaLM 2 later approached expert-level performance on several question-answering tasks, although its developers cautioned that benchmark scores do not establish safe performance in complex clinical practice (2). These studies established capability, but they did not resolve whether LLMs can gather information, manage ambiguity, apply guidelines, and make sequential decisions.
The potential applications extend across the clinical workflow. During diagnosis, an LLM might extract salient findings and generate a differential. It can also suggest tests and revise its assessment as new evidence emerges. For patient management, the model retrieves guidelines and identifies contraindications while comparing medical options and drafting treatment plans. Additional proposed uses include triage, documentation, surgical planning and patient communication. Reviews of clinical and surgical applications have described each of these functions while consistently concluding that LLMs should complement rather than replace clinicians (3). A scoping review of measured performance and clinician perceptions likewise found considerable interest in LLM-based decision support, but noted that enthusiasm has outpaced the maturity of objective clinical evidence (4).
The risks of clinical use differ in kind from errors in ordinary consumer applications. An LLM may generate a plausible falsehood, omit an urgent diagnosis, recommend an unsafe action, or answer semantically similar prompts inconsistently. It may also express confidence that is poorly calibrated to correctness or reproduce biases embedded in its training data. The evaluation literature remains difficult to compare because studies use different cases, prompts, outcomes, and grading schemes. Park et al. (5) identified substantial methodological heterogeneity and called for clinical utility frameworks that include real patient data, implementation barriers, and social determinants of health. Ho et al. (6) similarly argued that factuality, relevance, completeness, reasoning quality, safety, and bias should be reported separately rather than collapsed into a single accuracy score.
The central gap is therefore between medical knowledge and clinical decision-making. Static examinations mainly test recall and recognition. Patient care requires information gathering, prioritization, probabilistic reasoning, contextual judgment, and action under uncertainty. This review synthesizes evidence on LLM performance in differential diagnosis, diagnostic interviewing, test selection, treatment planning, triage, guideline adherence, and specialty-specific workflows. The relevant question is whether an LLM can contribute reliably to the sequence of decisions on which patient care depends.
Methods
The core evidence base of this structured narrative review includes 20 peer-reviewed sources that evaluated LLMs in clinical decision-making or synthesized clinically relevant performance (see Appendix A). PRISMA 2020 principles informed transparent reporting where applicable, including explicit eligibility criteria, organized extraction, and narrative synthesis. A complete multi-database search history, deduplication log, independent screening record, and PROSPERO registration number were not available.
Eligible evidence addressed diagnostic accuracy, differential diagnosis, diagnostic interviewing, clinical reasoning, test selection, treatment recommendations, guideline adherence, medication selection, triage, or specialty-specific decision support. The data sources ranged from real-world records and clinical vignettes to simulated encounters and comparisons with clinicians. Reviews were included when they synthesized objective clinical performance across models or specialties. Studies that lacked clinical decision performance outcomes were excluded from the core synthesis.
For each source, the review considered model identity and version, specialty, task, case format, comparator, outcome measure, and reported limitations. Accuracy was interpreted in relation to the stage of care being assessed. Selecting a final diagnosis after receiving complete information was treated as distinct from forming an early differential, selecting tests, conducting an interview, assigning triage acuity, or recommending management. Safety outcomes included hallucination, fabricated evidence, guideline deviation, inappropriate reassurance, overtriage or undertriage, bias, prompt sensitivity, and incorrect tool use.
Formal meta-analysis was not feasible due to limited homogeneity between model versions, prompts, grading standards and specialties. Methodological quality and risk of bias were assessed at the study level using design-appropriate criteria. Diagnostic accuracy and classification studies were evaluated with adapted QUADAS-2 domains (7), nonrandomized human-comparator studies were additionally informed by ROBINS-I principles (8), and systematic or scoping reviews were assessed using relevant AMSTAR 2 domains (9). The evaluation focused on three primary domains: care parameters, study design, and reporting integrity. Specific considerations included case selection and applicability, alongside technical variables like model-version transparency, prompt standardization, repeated sampling, and evaluator blinding. Furthermore, data sets were checked for missing data, selective reporting, and reference-standard validity. Ratings were classified as low, moderate, or high concern. Because the appraisal tools differ by design, the overall ratings are structured within-design judgments and should not be interpreted as directly equivalent across primary studies and evidence syntheses. Conceptual sources without comparative data were not assigned an overall rating.
Results
The core evidence base comprised eleven empirical studies and nine systematic, scoping, or narrative reviews. Across these designs, a consistent pattern emerged: accuracy depended more on the structure of the task than on any general notion of clinical competence. Models tended to perform better when the relevant facts were already assembled and the requested output was constrained. Reliability declined when the model had to identify what information mattered, seek missing evidence, choose among tools, or act safely under uncertainty. Comparative analyses also found meaningful differences among model families and generations, so results from one version should not be generalized indefinitely (10).
Broad clinical reasoning evaluations illustrate this distinction. Hager et al. (11) tested LLMs across realistic decision pathways derived from clinical records. Leading models did not consistently match clinician diagnostic accuracy, frequently departed from diagnostic and treatment guidelines, and sometimes performed worse when additional information was supplied. Their responses were sensitive to prompt wording and information order. Some models hallucinated tools that did not exist. Rao et al. (12) evaluated 21 LLMs across 29 standardized vignettes and five domains: differential diagnosis, diagnostic testing, final diagnosis, management, and miscellaneous reasoning tasks. Newer reasoning-oriented models improved on several measures, but early diagnostic reasoning remained the weakest stage. The authors concluded that off-the-shelf systems were not ready for unsupervised patient-facing use.
Advanced reasoning models nevertheless achieved strong results on selected physician tasks. Brodeur et al. (13) reported high performance across experiments in differential diagnosis, management reasoning, probabilistic reasoning, and emergency-department second opinions. The comparison used blinded model outputs and historical physician benchmarks, not autonomous real-time management. Its findings therefore demonstrate what a contemporary model can produce under controlled conditions, but they do not establish prospective reliability, safe tool use, or improved patient outcomes. The apparent contrast with Hager et al. (11) is best understood as task specificity. A model may generate an excellent differential from a curated note and still struggle when it must gather information, select tests, or follow an entire treatment pathway.
Conversational and multi-agent systems yielded some of the most encouraging findings. In simulated diagnostic encounters, AMIE collected information through text dialogue and produced differential diagnoses that specialist evaluators rated as more accurate and complete than those of primary care physicians in the study (14). It also performed well on communication measures. The encounter remained a simulated text exchange, however, without physical examination, nonverbal behavior, interruptions, or fragmented records. Chen et al. (15) tested a multi-agent framework in 302 rare-disease cases, using several doctor agents and a supervising agent to approximate multidisciplinary discussion. The framework outperformed single-agent GPT-3.5 and GPT-4 configurations. Structured deliberation may reduce some errors, but it remains dependent on the base model and does not by itself establish clinical safety.
Specialty-specific evidence was similarly uneven. A systematic review and meta-analysis of GPT-based radiology studies found substantial variation by model, prompt, input modality, and subspecialty (16). Many studies supplied textual case descriptions rather than raw images, which means the model was reasoning from an interpreted report rather than performing image analysis. In oncology, Ferber et al. (17) developed a GPT-4-based agent that coordinated radiology, segmentation, histopathology, molecular, and calculation tools. The study illustrated the potential value of retrieval and tool orchestration in a structured specialty, but also showed that less capable models could misuse tools, perform nonsensical calculations, or invent imaging findings.
Cardiology reviews reported relatively strong results for educational questions, patient information, and selected electrocardiographic or chronic-disease tasks, but weaker performance in acute guidance (17). In emergency triage, GPT-4 weighted age, vital signs, and risk differently from medical experts assigning Emergency Severity Index acuity, supporting cautious clinician-reviewed use rather than independent triage (19). Psychiatry poses related difficulties because risk assessment depends on context, ambiguity, rapport, and subtle language. A systematic review identified possible roles in diagnosis, education, and communication, but found limited real-patient validation and marked methodological variation (20). In youth mental health emergency vignettes, ChatGPT showed promise as a supportive triage aid, yet the small simulated sample and limited clinician comparison did not justify autonomous crisis assessment (21).
Safety studies identified vulnerabilities that conventional accuracy measures may overlook. Granstedt et al. (22) defined hallucination in medical devices as a plausible error whose clinical significance depends on its effect on the task. This framing matters because polished language can lend credibility to an incorrect recommendation. Omar et al. (23) inserted a single fabricated laboratory value, sign, or condition into 300 simulated clinical notes and tested six models. The systems repeated or elaborated on the planted falsehood in up to 83% of cases. A mitigation prompt reduced the rate but did not eliminate it. Other clinically relevant failures include invented citations, unsupported guideline claims, omitted contraindications, and confidence that is not justified by the evidence.
Taken together, the literature supports a supervised rather than autonomous model of use. Clinicians may use an LLM to broaden a differential, retrieve a possible guideline, summarize a long chart, or draft documentation, while retaining responsibility for verification and action. This arrangement is not automatically safe. Clinicians can anchor on an incorrect suggestion or defer to confident language. Even so, the available evidence indicates that reviewable outputs and meaningful clinician oversight provide safeguards that are absent when the model acts independently (3,4,11,12,14).
Table 1
Summary of Performance Patterns Across Clinical Decision Tasks
Note: The table summarizes recurrent patterns across heterogeneous designs and should not be interpreted as a pooled estimate.


Methodological Quality and Risk-of-Bias Assessment
The structured assessment identified no source with uniformly low overall concern after applicability was considered. Fourteen sources were rated as moderate concern and five as high concern. The conceptual article on hallucinations in medical devices was not formally rated because it did not report comparative clinical accuracy data. The most common concerns were reliance on simulated or publicly available cases, possible training-data contamination, small specialty samples, subjective or incompletely blinded grading, and limited prospective validation. High-concern ratings were concentrated in small proof-of-concept studies and reviews without a formal study-level appraisal.
Table 2
Study-Level Methodological Quality and Risk-of-Bias Assessment
Note: Ratings are within-design judgments informed by QUADAS-2, ROBINS-I, and AMSTAR 2. They are not directly equivalent across primary studies and reviews.


Discussion
Accuracy varies depending on the LLM model, the clinical task, and the information provided. A correct final diagnosis should, but often does not, demonstrate competence in the earlier stages of a workup. Therefore, current models can be highly capable under structured conditions, but their performance becomes less dependable when the conditions of evaluation change.
The distinction between knowledge and process is particularly important. Medical LLMs contain broad clinical information and can produce answers that clinicians judge useful (1,2). Clinical reasoning, however, unfolds over time. A physician decides what to ask, what to examine, what to test, how to revise probabilities, and when action is necessary before certainty is possible. Hager et al. (11) found that models could struggle to prioritize information as it accumulated, while Rao et al. (12) documented persistent weakness in the early stages of reasoning despite stronger performance in final diagnosis and management. Current LLMs may therefore be more reliable as synthesis tools than as independent clinical agents.
Structured systems offer a plausible route to improvement. Multi-agent deliberation, retrieval-augmented generation, explicit guideline grounding, and specialized tools can decompose a difficult task and create opportunities for verification. The multi-agent results suggest that critique and supervised consensus can improve diagnostic performance (15). The oncology agent similarly shows how an LLM can coordinate external tools rather than rely solely on generated text (17). Additional components also introduce new failure modes, including incorrect tool selection, propagation of an early error, retrieval of outdated guidance, and false consensus among agents that share the same underlying model. Evaluation should therefore address the performance of the full system, not only the language model at its center.
Initial clinical deployment should concentrate on tasks in which errors are detectable and outputs can be reviewed. Appropriate candidates include drafting documentation, organizing information, generating a provisional differential for clinician review, identifying unanswered questions, and retrieving guideline passages with verifiable sources. Autonomous triage, medication initiation, discharge decisions, and independent treatment planning require a much higher evidentiary threshold. In under-resourced settings, LLMs may eventually expand access to decision support, but limited staffing does not reduce the importance of safety, since local validation and workforce preparation become more consequential. Similarly, telehealth education for clinicians serving underserved communities has shown that technology alone does not ensure effective care without training and attention to implementation barriers (24).
Ethical and governance questions extend beyond aggregate accuracy. Responsibility may be divided among the clinician, health system, and model developer. Explanations may sound coherent without faithfully representing how the model reached its answer. Bias may appear through differential recommendations, uneven language performance, or lower accuracy for populations poorly represented in training data. Clinical information may also be exposed when it is transmitted to external systems. Because model updates can change behavior without an obvious change in the interface, validation can become obsolete quickly. These concerns reinforce the need for transparent reporting, ongoing monitoring, and explicit accountability structures (1,5,6).
The underlying literature has notable limitations. Many studies rely on published case reports, researcher-written vignettes, or static records that have already been curated by clinicians. These formats remove much of the interpersonal complexity of practice, and public cases may have appeared in the LLM’s training data. Model versions, prompts, sampling settings, and access dates are often incompletely described. Outcomes range from exact-match diagnosis to expert-rated usefulness, sometimes without blinded evaluation. Small samples remain common in specialty and triage studies. Reviews in radiology, cardiology, and psychiatry each noted marked heterogeneity and little prospective validation (16,18,20). Direct numerical comparisons should therefore be interpreted cautiously.
Future evaluations should use standardized, clinically meaningful methods. Technical reproducibility must be ensured by fully detailing the model identifier, access date, and prompt engineering strategies. Authors should also specify active sampling parameters, the number of experimental runs, and how the model handled system refusals or external tools. To test these models rigorously, datasets need to feature a balanced mix of common and rare clinical conditions. These sets should explicitly introduce real-world complications, such as incomplete information, distractors, or conflicting evidence. Furthermore, testing must challenge the model’s capacity for numerical calculations while accounting for demographic variations and time-sensitive emergencies. Evaluation metrics should avoid generalized scoring by measuring distinct clinical phases independently. Researchers should isolate upstream tasks like information gathering and test selection from downstream outcomes like differential diagnosis, final management, and triage. This granular approach ensures that critical safety metrics, like guideline adherence, hallucination rates, and potential patient harm, are tracked as separate data points. Moving forward, prospective studies should transition from these isolated benchmarks to real-world environments. These studies should compare usual care with clinician-plus-AI workflows and use patient outcomes and workflow consequences as primary endpoints.
Human oversight also requires direct study. Oversight can work when an interface makes uncertainty visible and links claims to sources, allowing clinicians to reject a recommendation easily. It can fail when time pressure, automation bias, or poorly designed alerts promote uncritical acceptance. Training should therefore address not only how to prompt an LLM, but how to audit its claims and recognize hallucinations. The most credible near-term model is a calibrated partnership in which the LLM expands informational capacity and the clinician remains accountable for the decision.
Conclusion
Large language models have progressed from medical question-answering systems to conversational multi-agent platforms capable of participating in complex clinical tasks. The literature documents meaningful strengths in medical knowledge, final-diagnosis selection, diagnostic interviewing, and documentation. It also identifies recurrent weakness in early diagnostic reasoning, test selection, acute triage, and guideline adherence under uncertainty.
The present evidence supports adjunctive clinical decision support rather than autonomous care. Clinicians can benefit from the breadth and speed of an LLM while supplying the contextual judgment and accountability that the technology lacks. Before higher-risk deployment, these systems require prospective real-world trials, standardized evaluation, specialty-specific validation, and effective hallucination controls. These systems must be re-evaluated following new model updates and should be integrated into workflow designs in a manner that preserves meaningful human control. Instead of strictly evaluating models for their ability to outperform physicians on selected benchmarks, future studies should further explore the ability of the human-AI system to reliably, equitably, and safely improve care.
Works Cited
-
Singhal, Karan, et al. “Large language models encode clinical knowledge.” Nature, vol. 620, no. 7972, 12 July 2023, pp. 172-180, https://doi.org/10.1038/s41586-023-06291-2.
-
Singhal, Karan, et al. “Toward expert-level medical question answering with large language models.” Nature Medicine, vol. 31, no. 3, 8 Jan. 2025, pp. 943-950, https://doi.org/10.1038/s41591-024-03423-7.
-
Pressman, Sophia M., et al. “Clinical and surgical applications of large language models: A systematic review.” Journal of Clinical Medicine, vol. 13, no. 11, 22 May 2024, p. 3041, https://doi.org/10.3390/jcm13113041.
-
Delourme, Solene, et al. “Measured performance and healthcare professional perception of large language models used as clinical decision support systems: A scoping review.” Studies in Health Technology and Informatics, 22 Aug. 2024, https://doi.org/10.3233/shti240543.
-
Park, Ye-Jean, et al. “Assessing the research landscape and clinical utility of large language models: A scoping review.” BMC Medical Informatics and Decision Making, vol. 24, no. 1, 12 Mar. 2024, https://doi.org/10.1186/s12911-024-02459-6.
-
Ho, Cindy N., et al. “Qualitative metrics from the biomedical literature for evaluating large language models in clinical decision-making: A narrative review.” BMC Medical Informatics and Decision Making, vol. 24, no. 1, 26 Nov. 2024, https://doi.org/10.1186/s12911-024-02757-z.
-
Whiting, Penny F., et al. "Quadas-2: A revised tool for the quality assessment of Diagnostic Accuracy Studies." Annals of Internal Medicine, vol. 155, no. 8, 18 Oct. 2011, pp. 529-536, https://doi.org/10.7326/0003-4819-155-8-201110180-00009.
-
Sterne, Jonathan AC, et al. "Robins-I: A tool for assessing risk of bias in non-randomised studies of interventions." BMJ, 12 Oct. 2016, p. i4919, https://doi.org/10.1136/bmj.i4919.
-
Shea, Beverley J, et al. "Amstar 2: A critical appraisal tool for systematic reviews that include randomised or non-randomized studies of healthcare interventions, or both." BMJ, 21 Sept. 2017, https://doi.org/10.1136/bmj.j4008.
-
Wang, Ling, et al. “Accuracy of large language models when answering clinical research questions: Systematic review and network meta-analysis.” Journal of Medical Internet Research, vol. 27, 30 Apr. 2025, https://doi.org/10.2196/64486.
-
Hager, Paul, et al. “Evaluation and mitigation of the limitations of large language models in clinical decision-making.” Nature Medicine, vol. 30, no. 9, 4 July 2024, pp. 2613-2622, https://doi.org/10.1038/s41591-024-03097-1.
-
Rao, Arya S., et al. “Large language model performance and clinical reasoning tasks.” JAMA Network Open, vol. 9, no. 4, 13 Apr. 2026, https://doi.org/10.1001/jamanetworkopen.2026.4003.
-
Brodeur, Peter G., et al. “Performance of a large language model on the reasoning tasks of a physician.” Science, vol. 392, no. 6797, 30 Apr. 2026, pp. 524-527, https://doi.org/10.1126/science.adz4433.
-
Tu, Tao, et al. “Towards Conversational Diagnostic Artificial Intelligence.” Nature, vol. 642, no. 8067, 9 Apr. 2025, pp. 442-450, https://doi.org/10.1038/s41586-025-08866-7.
-
Chen, Xi, et al. “Enhancing diagnostic capability with multi-agents conversational large language models.” Npj Digital Medicine, vol. 8, no. 1, 13 Mar. 2025, https://doi.org/10.1038/s41746-025-01550-0.
-
Nguyen, Daniel, et al. “A systematic review and meta-analysis of GPT-based differential diagnostic accuracy in radiological cases: 2023-2025.” Frontiers in Radiology, vol. 5, 28 Oct. 2025, https://doi.org/10.3389/fradi.2025.1670517.
-
Ferber, Dyke, et al. “Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology.” Nature Cancer, vol. 6, no. 8, 6 June 2025, pp. 1337-1349, https://doi.org/10.1038/s43018-025-00991-6.
-
Gendler, Moran, et al. “Large language models in Cardiology: Systematic Review.” JMIR Cardio, vol. 10, 16 Apr. 2026, https://doi.org/10.2196/76734.
-
Haim, Gal ben, et al. “Evaluating large language model-assisted emergency triage: A comparison of acuity assessments by gpt-4 and medical experts.” Journal of Clinical Nursing, 28 Nov. 2024, https://doi.org/10.1111/jocn.17490.
-
Omar, Mahmud, et al. “Applications of large language models in psychiatry: A systematic review.” Frontiers in Psychiatry, vol. 15, 24 June 2024, https://doi.org/10.3389/fpsyt.2024.1422807.
-
Thotapalli, Samanvith, et al. “Potential of chatgpt in Youth Mental Health Emergency Triage: Comparative analysis with clinicians.” Psychiatry and Clinical Neurosciences Reports, vol. 4, no. 3, 15 July 2025, https://doi.org/10.1002/pcn5.70159.
-
Granstedt, Jason, et al. “Hallucinations in medical devices.” Artificial Intelligence in the Life Sciences, vol. 8, Dec. 2025, p. 100145, https://doi.org/10.1016/j.ailsci.2025.100145.
-
Omar, Mahmud, et al. “Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support.” Communications Medicine, vol. 5, no. 1, 2 Aug. 2025, https://doi.org/10.1038/s43856-025–01021-3.
-
Dolouei, I., et al. “Preparing Medical Students with the Knowledge to Provide Telehealth in Underserved Communities.” Journal of Healthcare Solutions, vol 3, 2025, https://doi.org/10.58417/OWFD3072.
How to Cite: Bhullar, S. and A. Elhaija. "Accuracy and Safety of Large Language Models in Medical Diagnostics and Clinical Decision-Making: A Structured Narrative Review." Journal of Healthcare Solutions, vol. 4, no. 2, 2026.
Appendix A
Characteristics of the Core Evidence Base
Note: The table is descriptive. It does not imply direct comparability across model generations, tasks, specialties, or study designs

