The readability and quality paradox: Comparing ChatGPT, Gemini, and Perplexity outputs on pediatric chest pain queries.
پخش حرفهای فارسی و انگلیسی
در حال بررسی نسخههای صوتی ذخیرهشده…
تنظیم صدای طبیعی و سرعت
صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده میشود معمولاً طبیعیترند. انتخاب صدا به صداهای نصبشده در ویندوز و مرورگر شما بستگی دارد.
چکیده اصلی
Pediatric chest pain represents a frequently encountered symptom in emergency departments and outpatient clinics throughout childhood, often generating substantial parental anxiety. This study aims to comparatively analyze the readability, information quality, and scientific reliability of textual contents generated by the artificial intelligence (AI) chatbots ChatGPT, Gemini, and Perplexity regarding pediatric chest pain, utilizing multidimensional analytical indices. Out of the top 25 queries with the highest global search volume on Google Trends, 17 unique keywords meeting the predefined inclusion criteria were filtered on June 1, 2026. These inquiries were directed to all three AI platforms in distinct, independent user sessions, yielding a total of 51 textual responses. Linguistic readability levels were calculated across six separate digital interfaces using the FKGL, FRES and other formulas, and the results were benchmarked against the sixth-grade reading level" threshold. Scientific reliability was audited via the Modified DISCERN and JAMA benchmarks, while content quality was examined using the GQS and EQIP instruments. The median readability scores computed across all three AI models were found to be statistically and significantly above the targeted sixth-grade comprehension threshold, indicating a high level of difficulty (p < 0.001). Post-hoc pairwise evaluations revealed that ChatGPT and Gemini offered linguistically more accessible and structurally less complex textual architectures; conversely, Perplexity exhibited a significantly more difficult and academic linguistic framework (p < 0.0167). On the other hand, regarding content quality and source credibility, Perplexity demonstrated remarkably superior and more optimized scores across all evaluated instruments, namely GQS (p = 0.002, p = 0.001), JAMA (p < 0.001, p < 0.001), mDISCERN (p = 0.005, p = 0.001), and EQIP (p = 0.003, p < 0.001), compared directly to both ChatGPT and Gemini, respectively. Although generative AI formulations harbor considerable potential to provide extensive data on pediatric chest pain, the sophisticated, university-level linguistic architecture of these texts poses a substantial access barrier for individuals who lack proficient digital health literacy skills. While Perplexity achieved the performance closest to the "gold standard" in terms of informational accuracy, no model is currently mature enough to substitute for a professional medical consultation due to observed citation biases and omissions. It is of paramount importance that clinicians actively guide families to scrutinize online health data through a critical lens.
نتیجه فارسی
این مطالعه خروجیهای هوش مصنوعی در مورد دردهای قفسه سینه کودکان را بررسی کرد. Perplexity در کیفیت محتوا و منابع برتری داشت، اما خروجیها عموماً بسیار دشوار و سطح دانشگاهی بودند. هیچ مدلای هنوز جایگزین مشاوره پزشکی نیست.
- خروجیهای هوش مصنوعی در مورد دردهای قفسه سینه کودکان بسیار دشوار بودند.
- Perplexity در کیفیت محتوا و منابع برتری داشت.
- خروجیها سطح دانشگاهی داشتند و برای افراد با سواد سلامت دیجیتال پایین دشوار بودند.
- هیچ مدلای هنوز جایگزین مشاوره پزشکی حرفهای نیست.
- پزشکان باید خانوادهها را برای بررسی دادههای آنلاین با دقت راهنمایی کنند.
ترجمه فارسی چکیده
این مطالعه به مقایسه خوانایی، کیفیت اطلاعات و اعتبار علمی محتوای متنی تولید شده توسط هوش مصنوعی (ChatGPT، Gemini و Perplexity) در مورد دردهای قفسه سینه کودکان میپردازد. از میان 25 پرسش محبوبتر در گوگل تردز، 17 کلمه کلیدی فیلتر شد و به هر سه پلتفرم هوش مصنوعی پاسخ داده شد. میانگین امتیازات خوانایی بالاتر از سطح درک ششم دبستان بود (p < 0.001). Perplexity در ساختار زبانی و دستهبندی، دشوارتر از ChatGPT و Gemini بود (p < 0.0167)، اما در کیفیت محتوا و اعتبار منابع، امتیازات برتری نشان داد (p = 0.002، p = 0.001 و غیره). هرچند Perplexity به استاندارد طلایی نزدیکتر بود، اما به دلیل سوگیریها و حذف منابع، هنوز جایگزین مشاوره پزشکی حرفهای نیست.
روش پژوهش
از 25 پرسش محبوب در گوگل تردز استفاده شد. 17 کلمه کلیدی به هر سه پلتفرم هوش مصنوعی پاسخ داده شد. خوانایی با فرمولهای FKGL و FRES و کیفیت با ابزارهای GQS، JAMA و mDISCERN ارزیابی شد.
محدودیتها
محدودیتها در متن ارائه شده ذکر نشدهاند.
نمای PICO و پیامدها
- جمعیت
- کودکان مبتلا به درد قفسه سینه
- مداخله/مواجهه
- پرسشهای مرتبط با درد قفسه سینه کودکان
- مقایسه
- ChatGPT و Gemini
- حجم نمونه
- 51 پاسخ متنی
متن کامل اصلی
برای بررسی دسترسی کتابخانهای یا خرید، رکورد اصلی را باز کنید.
رفتن به منبع اصلی