Open-access Reliability and quality of information provided by artificial intelligence chatbots on post-contrast acute kidney injury: an evaluation of diagnostic, preventive, and treatment guidance

SUMMARY

OBJECTIVE:  The aim of this study was to evaluate the reliability and quality of information provided by artificial intelligence chatbots regarding the diagnosis, preventive methods, and treatment of contrast-associated acute kidney injury, while also discussing their benefits and drawbacks.

METHODS:  The most frequently asked questions regarding contrast-associated acute kidney injury on Google Trends between January 2022 and January 2024 were posed to four artificial intelligence chatbots: ChatGPT, Gemini, Copilot, and Perplexity. The responses were evaluated based on the DISCERN score, the Patient Education Materials Assessment Tool for Printable Materials score, the Web Resource Rating scale, the Coleman-Liau index, and a Likert scale.

RESULTS:  As per the DISCERN score, the quality of information provided by Perplexity received a rating of "good", while the quality of information acquired by ChatGPT, Gemini, and Copilot was scored as "average." Based on the Coleman-Liau index, the readability of the responses was greater than 11 for all artificial intelligence chatbots, suggesting a high level of complexity requiring a university-level education. Similarly, the understandability and applicability scores on the Patient Education Materials Assessment Tool for Printable Materials and the Web Resource Rating scale were low for all artificial intelligence programs. In consideration of the Likert score, all artificial intelligence chatbots received favorable ratings.

CONCLUSIONS:  While patients increasingly utilize artificial intelligence chatbots to acquire information about contrast-associated acute kidney injury, the readability and understandability of the information provided may be low.

KEYWORDS:
Artificial intelligence; Contrast media; Acute kidney injury

INTRODUCTION

Contrast-associated acute kidney injury (CA-AKI) is defined as the deterioration of renal function within three days following the administration of intravascular iodinated contrast material, with no other identifiable cause playing a role in the etiology1. With the advancement of diagnostic and interventional techniques and the increasing use of contrast agents in radiology, CA-AKI has become a significant cause of iatrogenic acute kidney injury1. The incidence of CA-AKI is more than 2% in the general population2.

In a meta-analysis study conducted in 2019, which reviewed studies on CA-AKI from various countries, 23 studies from Turkey were observed. The study, which involved 17,945 individuals, found that CA-AKI developed in 2,571 patients (14.7%), and the mortality rate in patients with CA-AKI was 14.6%. In this meta-analysis, the CA-AKI rates were 14.7% in the United States, 12.4% in China, 14.5% in Japan, and 12.9% in Italy3.

The pathophysiology of CA-AKI involves vasoconstriction, oxidative stress, and hypoxia4. To prevent CA-AKI, some recommended methods include pre-hydration before contrast-enhanced imaging, the use of N-acetylcysteine, and minimizing the contrast dose to the lowest possible level5.

Artificial intelligence chatbots (AICs), deep learning, and digital health terms have been successfully integrated into daily practice6. An AIC is a broad computer technology that enables machines to perform tasks and interpretations that normally require human intelligence. Tasks such as natural language understanding, image recognition, decision-making, and problem-solving are encompassed within AICs7. Chatbots utilize natural language processing to respond to user queries by matching them with the pre-compiled optimal responses of the system8.

The recent launch of ChatGPT by OpenAI, which integrates deep learning and generative pre-training transformer architecture-based language models, has significantly enhanced the capabilities of chatbots. Information obtained online can be inconsistent, making it difficult to distinguish accurate information from misinformation9,10. Despite much of the information being commercially driven, patients may not discern this, and without appropriate guidance, internet-based information can be harmful, confusing, and misleading9.

The aim of this study was to evaluate the reliability and quality of information provided by AICs concerning the definition, preventive measures, and treatment of CA-AKI. Moreover, the study sought to address the benefits and drawbacks of using AICs from the users’ perspective in this context.

METHODS

Between January 2022 and January 2024, the most frequently asked questions related to CA-AKI on Google Trends were posed to four AICs: ChatGPT, Gemini, Copilot, and Perplexity. These questions included "What is CA-AKI?," "What are the preventive methods for CA-AKI?," and "What treatments are used for CA-AKI?." The responses provided by the AICs were evaluated blindly by two different radiologists, and the DISCERN score was calculated separately. The results were evaluated by two radiologists, and a single result was reported. The responses regarding CA-AKI were evaluated in reference to the European Society of Urogenital Radiology guidelines updated in 201811. DISCERN is a scoring-based questionnaire system that evaluates the quality of information provided to patients based on evidence and unbiased information. The DISCERN system consists of 16 questions scored on a scale of 1 to 5. The adequacy of the information is scored by the user based on the accuracy of the information. DISCERN scores are interpreted as follows: scores ranging from 16 to 26 points indicate poor quality, 27 to 38 points indicate low quality, 39 to 50 points indicate average quality, 51 to 62 points indicate good quality, and 63 to 75 points indicate excellent quality12,13.

The Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P) is a reliable and valid assessment system for evaluating the understandability and actionability of information obtained from printed brochures, flyers, websites, and audio and visual materials, for both individuals without medical education and healthcare professionals. Scores obtained for understandability and actionability are presented in percentages14.

The Web Resource Rating (WRR) scale, scored from 0 to 100%, is used to measure the reliability and transparency of information obtained online15. The Coleman-Liau index is employed to measure readability, and scores above 11 points indicate that reading is very difficult and suggest that at least a university-level education is required16.

In this study, median values were utilized for the scores obtained from the scoring systems, number of recommended treatments, and consistency of treatment options, while mean values were used to present the word count and Coleman-Liau index results.

Data availability

The data associated with the study are not publicly available but are available from the corresponding author on reasonable request.

RESULTS

The most frequently asked three questions related to CA-AKI were posed to four AICs by two radiologists. The scores acquired were interpreted, and different scores were re-evaluated to obtain a unified score. The responses regarding CA-AKI were evaluated in reference to the European Society of Urogenital Radiology guidelines updated in 201817.

Table 1 presents the data regarding the word count, Coleman-Liau index, PEMAT-P scores, WRR score, number of treatment options presented to the patient, and adherence to the guidelines. As per the Coleman-Liau index, the readability of the responses was above 11 in all AICs, indicating a complexity requiring a university-level education. The PEMAT-P understandability and actionability scores were low for all AICs. Similarly, the WRR scores for all AICs were low. Based on the Likert scale, the AICs performed very well in providing information about preventive and therapeutic procedures. Almost all the treatment options recommended by the guidelines were available in the AICs. Table 2 shows the DISCERN scores of the AICs, indicating the quality of information regarding disease identification and treatment.

Table 1
Evaluation of the results of readability, understandability, and reliability of the artificial intelligence chatbots.
Table 2
Median DISCERN scores of the artificial intelligence chatbots.

All AICs provided sufficient responses regarding the definition of CA-AKI, listing underlying causes, identifying those more commonly affected, and determining risk factors. Since ChatGPT did not cite sources for the information provided, making it difficult to assess the reliability of the information, it did not receive a high score in this aspect. However, Copilot and Perplexity excelled, particularly in referencing, earning high scores in this regard. All AICs provided adequate information on preventive methods and measures to prevent CA-AKI. Only Gemini directed users to healthcare professionals in the presence of any doubts about potential risks.

According to the DISCERN score, Perplexity provided information of good quality with a total score of 53 regarding CA-AKI, while the quality of information provided by ChatGPT with a score of 46, Gemini with a score of 44, and Copilot with a score of 48 was "average." The content presented by Perplexity was more easily understandable due to its concise presentation of the necessary information with the least word count. As all AICs contained general information, they did not offer personalized or specific treatment options. Additionally, none of the AICs addressed potential problems if individuals did not receive treatment or difficulties and complications that may arise during treatment. According to the DISCERN scoring, Perplexity received the highest score, followed by Copilot, ChatGPT, and Gemini.

DISCUSSION

Artificial intelligence (AI) was first used in medicine in the early 1950s18. AI has played a vital role worldwide in the last few decades19. Particularly, with the free usage of ChatGPT version 3.5, the use of AICs has become ubiquitous in daily life. Undoubtedly, the most significant area where AICs have rapidly grown is the field of medicine, where they have begun to be frequently used by patients to promptly access information, increasing their understanding of treatment methods and diagnoses pertaining to their conditions20.

Patients have long had a desire to obtain information about their diseases. A study conducted by Keten and Erkan revealed that patients’ use of YouTube videos for information regarding undescended testicles was very high; however, the reliability of the information was low, at around 26%21. YouTube videos rely entirely on the information provided by the content creator and are subjective.

In recent years, AI technologies have begun to compile internet-based medical information to provide more reliable and academic literature. Some AICs even support their academic information with references. AI systems provide training on a specific dataset to predict better outcomes and assist in solving complex problems with high reliability22. This has led to speculations being made regarding the potential replacement of healthcare professionals by AICs with the introduction of AI systems in the medical field23.

The growing population and longer life expectancy are increasing the need for healthcare professionals in society, which appears to make it impossible for professionals to meet every societal need. In this regard, the recommendation of the World Health Organization on "Self-Care Interventions for Health" stands out. Due to increasing health problems, it will not be possible for healthcare professionals to provide one-on-one care for every patient. In such cases, AICs will play an important role as assistants rather than replacing healthcare professionals24.

All AICs provided sufficient information on the definition of CA-AKI, listed the underlying causes in bullet points, identified the groups more frequently affected, and determined the risk factors. However, all AICs failed to provide adequate details on topics included in the DISCERN score, such as information on resources, possible outcomes if the patient declines treatment, impacts of treatment options on the patient's quality of life, and the risks of treatment. Consequently, they received relatively poor scores in the scoring system. Since the requests by patients to AICs are generally superficial, the responses are brief and concise without addressing every disease comprehensively. While the general information may be sufficient, they might fall short on specific topics, resulting in lower scores.

Copilot and Perplexity facilitated the referral of patients to relevant articles by providing information and referencing the respective sources. However, scholars in medicine and the humanities have long emphasized the value of personal narratives that provide perspective on patients’ lives. Numerous studies have been conducted on the value of listening to, reading, and writing patients’ disease narratives25.

Once patients acquire information about their diseases, they feel the need to learn about personalized treatment options in order to understand the potential challenges of treatment and be aware of the situations that they and their families may face during the treatment process. This indicates that while patients can obtain general information about their diseases from AICs, they still need healthcare professionals for more specific treatment options and recommendations.

In this study, the AICs provided sufficient information regarding CA-AKI and explained the underlying conditions according to the DISCERN score, categorizing the information by headings to facilitate understandability. However, all AICs fell short in personalizing information and addressing potential experiences during the treatment implementation process. Various cancer types rank at the top of search engine queries26. In diseases such as cancer that are stage-dependent, it is more challenging for AICs to provide personalized information. However, CA-AKI is a field independent of disease stages, containing more general information. As a result, the DISCERN scores were fair for Copilot, ChatGPT, and Gemini and good for Perplexity. The PEMAT-P scores, encompassing understandability and actionability, were high, indicating that the understandability and actionability of the information provided would be difficult for those without a university-level education.

In a urological study examining AICs conducted by Talyshinskii et al., AICs have become focal points in fields such as urological symptom control, health screening, patient education, counseling, lifestyle changes, conservative management, clinical decision support, and post-treatment follow-up care27. In a previous study by Sheikh et al. comprising questions about acute kidney failure and renal replacement therapies asked to AICs, ChatGPT generated responses with a high accuracy rate of 97% related to acute kidney failure. It demonstrated that ChatGPT 4.0 is consistent in interpreting and responding to acute kidney failure regardless of linguistic changes, and it has a high accuracy rate. Moreover, it emphasized that further research is necessary to investigate the direct impact of AI-generated responses on patient understanding and education outcomes28. In light of previous studies, AICs provide accurate and reliable information on acute kidney injury, but a sufficient educational background is required for understanding. Nevertheless, there is a lack of studies comparing this information across different AICs.

The essential limitation of this study is the subjective nature of scoring the information provided by the AICs, which may vary among evaluators. This limitation can be mitigated by having AICs themselves perform the scoring, potentially yielding more objective results, instead of relying on separate assessments by two radiologists.

CONCLUSION

With increasing life expectancy and the growing significance of preventive medicine, the availability of radiological diagnostic methods has increased in contemporary medical practice. Patients utilize AICs to acquire information related to the potential CA-AKI associated with contrast agent use. However, the readability and understandability of the information require a university-level education for comprehension. Therefore, AICs require further development to cater to the needs of society and individuals without a medical background.

  • Funding:
    none.

REFERENCES

  • 1 Mohammad HA, Eric C, James C, Chirag P, Huzaif Q, Vikas S, et al. Contrast induced nephropathy: pathophysiology, risk factors, and preventation. Saudi J Kidney Dist T. 2018;29:1-9. https://doi.org/10.4103/1319-2442.225199
    » https://doi.org/10.4103/1319-2442.225199
  • 2 Rudnick MR, Leonberg-Yoo AK, Litt HI, Cohen RM, Hilton S, Reese PP. The controversy of contrast-induced nephropathy with intravenous contrast: what is the risk? Am J Kidney Dis. 2020;75(1):105-13. https://doi.org/10.1053/j.ajkd.2019.05.022
    » https://doi.org/10.1053/j.ajkd.2019.05.022
  • 3 Lun Z, Liu L, Chen G, Ying M, Liu J, Wang B, et al. The global incidence and mortality of contrast-associated acute kidney injury following coronary angiography: a meta-analysis of 1.2 million patients. J Nephrol. 2021;34(5):1479-89. https://doi.org/10.1007/s40620-021-01021-1
    » https://doi.org/10.1007/s40620-021-01021-1
  • 4 Shams E, Mayrovitz HN. Contrast-induced nephropathy: a review of mechanisms and risks. Cureus. 2021;13(5):e14842. https://doi.org/10.7759/cureus.14842
    » https://doi.org/10.7759/cureus.14842
  • 5 Kusirisin P, Chattipakorn SC, Chattipakorn N. Contrast-induced nephropathy and oxidative stress: mechanistic insights for better interventional approaches. J Transl Med. 2020;18(1):400. https://doi.org/10.1186/s12967-020-02574-8
    » https://doi.org/10.1186/s12967-020-02574-8
  • 6 Nadkarni GN. Introduction to artificial intelligence and machine learning in nephrology. Clin J Am Soc Nephrol. 2023;18(3):392-3. https://doi.org/10.2215/CJN.0000000000000068
    » https://doi.org/10.2215/CJN.0000000000000068
  • 7 Trivedi A, Kaur EK, Choudhary C, Kunal, Barnwai P. Should AI technologies replace the human Jobs? In: Proceedings of the 2023 2nd international conference for innovation in technology (INOCON), Bangalore, India; 2023. https://doi.org/10.1109/INOCON57975.2023.10101202
    » https://doi.org/10.1109/INOCON57975.2023.10101202
  • 8 Dwivedi YK, Kshetri N, Hughes L, Slade EL, Jeyaraj A, Kar AK, et al. "Opinion paper: so what if chatGPT wrote it?" multidisciplinary perspectives on opportunities, challenges and implications of generative conversational AI for research, practice and policy. IJIM. 2023;71:102642. https://doi.org/10.1016/j.ijinfomgt.2023.102642
    » https://doi.org/10.1016/j.ijinfomgt.2023.102642
  • 9 Diaz JA, Griffith RA, Ng JJ, Reinert SE, Friedmann PD, Moulton AW. Patients’ use of the Internet for medical information. J Gen Intern Med. 2002;17(3):180-5. https://doi.org/10.1046/j.1525-1497.2002.10603.x
    » https://doi.org/10.1046/j.1525-1497.2002.10603.x
  • 10 Boer MJ, Versteegen GJ, Wijhe M. Patients’ use of the Internet for pain-related medical information. Patient Educ Couns. 2007;68(1):86-97. https://doi.org/10.1016/j.pec.2007.05.012
    » https://doi.org/10.1016/j.pec.2007.05.012
  • 11 European Society of Urogenital Radiology. The Contrast Media Safety Committee. ESUR Guidelines on contrast agents, European Society of Urogenital Radiology Version 10. 2018. Available from: https://www.esur.org/wp-content/uploads/2022/03/ESUR-Guidelines-10_0-Final-Version.pdf
    » https://www.esur.org/wp-content/uploads/2022/03/ESUR-Guidelines-10_0-Final-Version.pdf
  • 12 Charmock D. The DISCERN handbook: quality criteria for consumer health information on treatment choices. Radcliffe: University of Oxford and The British Library; 1998. p. 7-51.
  • 13 Rao EM, Smith GP. DISCERN scores of morphea information on YouTube. J Dermatolog Treat. 2022;33(5):2702-4. https://doi.org/10.1080/09546634.2022.2062278
    » https://doi.org/10.1080/09546634.2022.2062278
  • 14 Shoemaker SJ, Wolf MS, Brach C. Development of the Patient Education Materials Assessment Tool (PEMAT): a new measure of understandability and actionability for print and audiovisual patient information. Patient Educ Couns. 2014;96(3):395-403. https://doi.org/10.1016/j.pec.2014.05.027
    » https://doi.org/10.1016/j.pec.2014.05.027
  • 15 Dobbins M, Watson S, Read K, Graham K, Yousefi Nooraie R, Levinson AJ. A tool that assesses the evidence, transparency, and usability of online health information: development and reliability assessment. JMIR Aging. 2018;1(1):e3. https://doi.org/10.2196/aging.9216
    » https://doi.org/10.2196/aging.9216
  • 16 Coleman R. A data science approach to readability. IEEE Xplore; 2021. 2021:268-272.60:283. https://doi.org/10.1109/CSCI54926.2021.00116
    » https://doi.org/10.1109/CSCI54926.2021.00116
  • 17 Stacul F, Molen AJ, Reimer P, Webb JA, Thomsen HS, Morcos SK, et al. Contrast induced nephropathy: updated ESUR Contrast Media Safety Committee guidelines. Eur Radiol. 2011;21(12):2527-41. https://doi.org/10.1007/s00330-011-2225-0
    » https://doi.org/10.1007/s00330-011-2225-0
  • 18 Secinaro S, Calandra D, Secinaro A, Muthurangu V, Biancone P. The role of artificial intelligence in healthcare: a structured literature review. BMC Med Inform Decis Mak. 2021;21(1):125. https://doi.org/10.1186/s12911-021-01488-9
    » https://doi.org/10.1186/s12911-021-01488-9
  • 19 Khanna D. Use of artificial intelligence in healthcare and medicine. IJIERT. 2028; 5:2025.
  • 20 Jiang F, Jiang Y, Zhi H, Dong Y, Li H, Ma S, et al. Artificial intelligence in healthcare: past, present and future. Stroke Vasc Neurol. 2017;2(4):230-43. https://doi.org/10.1136/svn-2017-000101
    » https://doi.org/10.1136/svn-2017-000101
  • 21 Keten T, Erkan A. An investigation of the reliability of YouTube videos on undescended testis. J Pediatr Urol. 2022;18(4):515.e1-6. https://doi.org/10.1016/j.jpurol.2022.04.021
    » https://doi.org/10.1016/j.jpurol.2022.04.021
  • 22 Richardson JP, Smith C, Curtis S, Watson S, Zhu X, Barry B, et al. Patient apprehensions about the use of artificial intelligence in healthcare. NPJ Digit Med. 2021;4(1):140. https://doi.org/10.1038/s41746-021-00509-1
    » https://doi.org/10.1038/s41746-021-00509-1
  • 23 Ostherr K. Artificial intelligence and medical humanities. J Med Humanit. 2022;43(2):211-32. https://doi.org/10.1007/s10912-020-09636-4
    » https://doi.org/10.1007/s10912-020-09636-4
  • 24 Bodenheimer T, Lorig K, Holman H, Grumbach K. Patient self-management of chronic disease in primary care. JAMA. 2002;288(19):2469-75. https://doi.org/10.1001/jama.288.19.2469
    » https://doi.org/10.1001/jama.288.19.2469
  • 25 Moore AR. The missing medical text: humane patient care. Melbourne: Melbourne University Press; 1978.
  • 26 Hopkins AM, Logan JM, Kichenadasse G, Sorich MJ. Artificial intelligence chatbots will revolutionize how cancer patients access information: ChatGPT represents a paradigm-shift. JNCI Cancer Spectr. 2023;7(2):pkad010. https://doi.org/10.1093/jncics/pkad010
    » https://doi.org/10.1093/jncics/pkad010
  • 27 Talyshinskii A, Naik N, Hameed BMZ, Juliebø-Jones P, Somani BK. Potential of AI-driven chatbots in urology: revolutionizing patient care through artificial intelligence. Curr Urol Rep. 2024;25(1):9-18. https://doi.org/10.1007/s11934-023-01184-3
    » https://doi.org/10.1007/s11934-023-01184-3
  • 28 Sheikh MS, Thongprayoon C, Suppadungsuk S, Miao J, Qureshi F, Kashani K, et al. Evaluating ChatGPT's accuracy in responding to patient education questions on acute kidney injury and continuous renal replacement therapy. Blood Purif. 2024;53(9):725-31. https://doi.org/10.1159/000539065
    » https://doi.org/10.1159/000539065

Publication Dates

  • Publication in this collection
    02 Dec 2024
  • Date of issue
    2024

History

  • Received
    23 June 2024
  • Accepted
    18 Aug 2024
location_on
Associação Médica Brasileira R. São Carlos do Pinhal, 324, 01333-903 São Paulo SP - Brazil, Tel: +55 11 3178-6800, Fax: +55 11 3178-6816 - São Paulo - SP - Brazil
E-mail: ramb@amb.org.br
rss_feed Acompanhe os números deste periódico no seu leitor de RSS
Ir para o topo Reportar erro