Understanding Clinician Perceptions of GenAI: A Mixed Methods Analysis of Clinical Documentation Tasks

Our mixed-methods study provides critical insights into clinician perspectives on GenAI integration in clinical documentation, revealing a nuanced landscape of cautious optimism tempered by legitimate concerns about safety, control, and clinical autonomy. The clear preference for medium automation levels across all documentation tasks represents a pivotal finding that challenges assumptions about the optimal degree of AI assistance in healthcare.

This paper builds upon earlier works exploring the attitudes of primary care doctors towards text automation [9]. Our findings align with previous research, indicating that clinicians are generally receptive to implementing GenAI to streamline their workflow and improve efficiency. By reducing the administrative burden and saving time on documentation tasks, AI may provide an opportunity for doctors to reconnect with their patients by allowing more time for direct patient interaction and care [29], simultaneously enhancing the quality and efficiency of clinical documentation [30].

Recent work demonstrated initial deployment of GenAI into clinical settings. For example, ambient AI was tasked with clinical note generation achieving initial, promising results [7]; GenAI chatbot showed clinical reasoning capability on par with clinicians [31]; or, GenAI deployed to evaluate stroke management adherence to guidelines reaching agreement levels of experts [32]. These studies show the potential to integrate GenAI in tasks going beyond clinical documentation and pave the way to a stream of research on improving other elements of the clinical encounter.

Theoretical Implications for Human-AI Collaboration

The consistent preference for medium automation aligns remarkably with established models of human-automation interaction. Parasuraman et al.’s [25] framework suggests that intermediate automation levels often provide optimal balance between workload reduction and maintenance of situation awareness. Our findings extend this framework to the clinical documentation context, where the stakes of maintaining awareness are particularly high.

The preference distribution—42% for medium automation in Information Extraction, 32% in Summarization, and 31% in Speech-to-Text—reflects what we term “calibrated trust” in AI systems. This pattern suggests clinicians seek a collaborative relationship with AI rather than replacement or minimal assistance. The significantly higher safety ratings for medium automation (typically rated “Safe with Caution” by 70-80% of participants) compared to high automation (rated “Probably Unsafe” by 30-35%) provides empirical support for this interpretation.

Our findings also contribute to understanding the “automation paradox” in healthcare: while participants acknowledged that higher automation could maximize efficiency, they simultaneously recognized that it might compromise their ability to maintain clinical oversight. This sophisticated understanding challenges simplistic narratives about resistance to technology and instead reveals thoughtful consideration of human-AI collaboration dynamics.

Addressing the Implementation Challenge

Our findings highlight several critical requirements for successful GenAI implementation in clinical settings:

Flexible Automation Architecture

The strong preference for adjustable automation levels suggests that one-size-fits-all approaches will likely fail. Successful systems must allow clinicians to dynamically adjust automation levels based on case complexity, time constraints, and personal comfort. This requirement challenges current AI system design paradigms that typically offer fixed levels of assistance. In this context, flexible automation and moderate levels of oversight are crucial for controlling and mitigating the risk of biases [19].

Transparent Clinical Oversight

Participants’ emphasis on maintaining clinical control reflects not just personal preference but professional responsibility. Systems must provide clear mechanisms for healthcare professionals to understand key aspects of AI operations—such as the rationale behind specific outputs, the underlying data sources, and the algorithm’s decision-making logic—rather than requiring complete understanding of all AI decisions. This transparency requirement extends beyond simple explainability to include practical tools for clinical oversight that allow clinicians to validate and modify AI-generated content while maintaining ultimate responsibility for documentation accuracy.

Robust Testing and Validation

The repeated emphasis on adequate trialing before deployment reflects clinicians’ understanding of the stakes involved in clinical documentation. Participants wanted evidence of system performance in real-world clinical settings, not just laboratory benchmarks. This suggests that implementation strategies must include extensive pilot testing with clear metrics for safety and effectiveness. GenAI systems must be calibrated to handle diverse populations and account for local nuances, as well as mitigate biases in training data [9, 15].

Addressing Medico-Legal Concerns

The emergence of liability and consent issues as major themes indicates that technical solutions alone are insufficient. Successful implementation requires clear policy frameworks addressing responsibility for AI-generated content, patient consent for voice recording, and integration with existing medico-legal structures. These frameworks must be developed collaboratively with clinical, legal, and regulatory stakeholders.

Integration with Patient-Centered Care

An important consideration for GenAI implementation is how it facilitates patient access to their own health records. In Australia, the My Health Record system provides patients with digital access to their health information, including discharge summaries, specialist letters, and clinical documents uploaded by healthcare providers [33]. With over 90% of Australians having a My Health Record [33], clinicians are increasingly aware that patients can view the documentation they create. This transparency adds further challenges on top of clinicians’ perspectives on AI-assisted documentation as the user and consumer of these automated clinical notes is no longer clinicians only, but also patients which may have a different set of requirements and views on what should be recorded in EHRs

Implications for GenAI Design and Deployment

Our findings suggest four design principles for clinical GenAI systems, as outlined in Table 6.

Table 6 Design principles for clinical GenAI systemsStrengths and Limitations

Our study has several strengths. To the best of our knowledge, this is the first work studying the implementation of LLM-based approaches in EHR, exploring multiple text-processing tasks and automation levels. These addressed day-to-day problems faced by primary care doctors in their practice, going beyond synthetic NLP benchmarks and hypothetical use cases. Furthermore, the qualitative clinician input, coded thematically and intertwined with quantitative analyses, offers invaluable insight for future research and practical EHR management system development. The sample size of 38 participants aligns with established guidelines for quantitative usability studies in HCI [34, 35], which recommend 40 participants for achieving a 15% margin of error with 95% confidence [35].

Several limitations warrant consideration when interpreting our findings:

First, the relatively small sample size, albeit sufficient for usability studies, and focus on Australian primary care physicians may limit the generalizability of our findings. However, it should be noted that the Australian healthcare system combines elements of both UK-style GP-based care and American-style healthcare delivery with private specialists and reimbursement schemes similar to Medicare, potentially making these insights relevant to multiple healthcare contexts. Nevertheless, rural practitioners, specialists, and clinicians in other healthcare systems may have different perspectives and requirements. The described tasks however are fruit of primary care knowledge and experiences that are likely applicable to multiple contexts and health systems (reviewing letters from specialists, transcribing conversations or summarizing content from the health record, thus our findings are likely to resonate for many clinicians, health systems and different environments.

Second, our usability study using prototype-based evaluation and synthetic patient data represents an essential preliminary step in the health technology development pipeline. While synthetic data cannot capture all nuances of real patient cases, our validation showed that participants rated the scenarios as highly representative of their day-to-day practice (mean ratings above 5 on a 7-point scale across all tasks), suggesting the synthetic cases successfully captured authentic clinical documentation challenges. This approach follows established HCI best practices for understanding user requirements before system implementation. This foundational usability research is a prerequisite for developing systems that will be both acceptable to clinicians and effective in practice.

Third, the cross-sectional design captures only initial impressions. Longitudinal research is essential to understand how perceptions and usage patterns evolve with extended exposure to GenAI systems.

Fourth, our study focused on subjective perceptions rather than objective outcomes. Future research should examine actual documentation quality, time savings, and clinical outcomes associated with different automation levels.

Finally, while our mixed-methods approach provides rich insights, the study would benefit from additional objective measures of usability and performance in actual clinical settings.

Comments (0)

No login
gif