From Subjectivity to Precision: Removing Inconsistency in Medical Evaluations
Written by Leonardo Viola and Francesco Rocchi
Current scientific literature is unequivocal: inconsistency in the interpretation of medical images and the consequent interobserver variability constitute a global challenge across various specialties. It is estimated that diagnostic errors affect between 3% and 5% of all diagnoses worldwide, which can result in up to 40 million errors annually, many of them linked to failures in the interpretation of imaging exams.
This interpretation is described as an inherently subjective and operator-dependent process, especially in modalities such as ultrasound. A study on thyroid nodules revealed that agreement among experienced radiologists is only fair for crucial characteristics such as echogenicity, shape, and margins. Although the reproducibility of reports was considered adequate, the ability to make a final diagnostic assessment was inconsistent.
In the case of head and neck modalities, in the assessment of extranodal extension (iENE), interobserver agreement was moderate for binary classifications (present/absent), but weak for severity grading. Surprisingly, even in binary assessments, 15% of the examinations were classified differently by 20% or more of the specialist radiologists.

Our advanced technology transforms multi-modality medical exams into intelligent digital entities
Causes of Inconsistencies
Analysis of the state of the art identifies a lack of standardization as one of the main factors contributing to diagnostic discrepancies, particularly the absence of unified global criteria. A clear example of this variability occurs in carotid ultrasound: in the US alone, more than 60 different peak systolic velocity thresholds are used to classify stenoses. Additionally, heterogeneity and lack of calibration among medical equipment directly compromise the reproducibility of results.
Alongside this technological inconsistency, the clinical reading process itself introduces a significant margin of error and arbitrariness. As already mentioned, the interpretation of medical images is an inherently subjective process, conditioned by the sensory perception and cognitive biases of each radiologist. This interobserver variability is accentuated in the absence of systematic training or calibration tools, such as educational atlases, making reports highly dependent on personal experience and professional practice time in the respective specialty.
Diagnostic consistency also faces two major practical challenges: workload overload, which generates fatigue and potentiates human error, and the intrinsic limitations of certain modalities. Ultrasound, for example, stands out as being strongly ‘operator-dependent’, linking the accuracy of the diagnosis to the technical skill of the person performing the examination. Furthermore, the structural complexity and subtlety of borderline lesions (such as the definition of ‘clustered nodules’ or the distinction between invasion and mere muscle displacement) generate frequent disagreements, in which minimal tissue changes may be interpreted as pathological by one physician and as normal variants by another.

QLens emerges as an innovative tool that mitigates interobserver variability in biomedical analysis
Clinical Consequences
This interpretative variability transcends mere statistical analysis, directly impacting the patient’s clinical management. Discrepancies in diagnostic criteria can result in incorrect classification of the severity of the pathology, leading to both inappropriate surgical referrals and omission of necessary interventions.
Consequently, patients with the same clinical picture face divergent therapeutic guidelines depending on the health unit or the physician consulted. Furthermore, the unreliability of tests triggers redundant additional investigations, such as biopsies or other unnecessary medical exams, overburdening health systems.

The QLens Software ensures the comparability and longitudinal continuity of studies over time
Our Solution
QLens technology emerges as a real and structural solution to mitigate interobserver variability. By converting clinical examinations into structured and measurable cognitive datasets, this system replaces interpretive subjectivity with objective radiomics metrics. Through the extraction of numerical data that is independent of the operator’s individual perception, QLens resolves the low agreement between specialists in the evaluation of qualitative criteria, such as echogenicity or nodule margins.
Supported by deep learning algorithms, this technology achieves millimeter precision in automatic segmentation and the identification of subtle biomarkers, overcoming the limitations of suboptimal image quality. Additionally, through Explainable Artificial Intelligence, QLens generates heat maps (like Grad-CAM) that correlate algorithmic decisions with specific visual evidence. This mechanism acts as a cognitive enhancer and an objective “third opinion” for the clinician, aligning machine reasoning with human judgment and drastically reducing errors in disease severity classification.
In addition to standardizing diagnosis in the present, QLens ensures the comparability and longitudinal continuity of studies over time, making examinations from different institutions or equipment technically comparable.
In short, by preserving the history of reviews and mapping global pathological patterns, QLens establishes a robust institutional memory that transforms the medical report into a document based on verifiable quantitative evidence, ensuring equity in the treatment of each patient.
References
Abdulaal, L., Maiter, A., Salehi, M., Sharkey, M., Alnasser, T., Garg, P., Rajaram, S., Hill, C., Johns, C., Rothman, A. M. K., Dwivedi, K., Kiely, D. G., Alabed, S., & Swift, A. J. (2024). A systematic review of artificial intelligence tools for chronic pulmonary embolism on CT pulmonary angiography. In Frontiers in Radiology (Vol. 4). Frontiers Media SA. https://doi.org/10.3389/fradi.2024.1335349
Abou-Foul, A. K., Kristunas, C., Henson, C., Andrew, D., Bidault, F., Dankbaar, J. W., Gebrim, E. S., de Graaf, P., Kim, J., King, A. D., Kuno, H., Medrano-Martorell, S., Goh, J. P. N., Santos-Armentia, E., Schafigh, D. G., Wu, X. C., Xiao, Y., Huang, S. H., Lydiatt, W. M., … Mehanna, H. (2026). Correlation between head and neck radiologists reporting on extranodal extension detected on radiological imaging: A head and neck cancer international group multinational study. Oral Oncology, 178. https://doi.org/10.1016/j.oraloncology.2026.107980
Alyami, J., Almutairi, F. F., Aldoassary, S., Albeshry, A., Almontashri, A., Abounassif, M., & Alamri, M. (2022). Interobserver variability in ultrasound assessment of thyroid nodules. Medicine (United States), 101(41), E31106. https://doi.org/10.1097/MD.0000000000031106
Du, K., Nair, A. R., Shah, S., Gadari, A., Vupparaboina, S. C., Bollepalli, S. C., Sutharahan, S., Sahel, J. A., Jana, S., Chhablani, J., & Vupparaboina, K. K. (2024). Detection of Disease Features on Retinal OCT Scans Using RETFound. Bioengineering, 11(12). https://doi.org/10.3390/bioengineering11121186
Fotopoulos, D., Ladakis, I., Filos, D., Moreno-Sánchez, P. A., van Gils, M., & Chouvarda, I. (2026). Explainable AI in Cancer Imaging: Scoping Review of Methods, Modalities, and Clinical Integration. In Journal of medical Internet research (Vol. 28, p. e80645). https://doi.org/10.2196/80645
Leoncini, A., & Trimboli, P. (2025). Interobserver agreement between artificial intelligence models in the thyroid imaging and reporting data system (TIRADS) assessment of thyroid nodules. Endocrine, 89(1), 197–201. https://doi.org/10.1007/s12020-025-04272-1
Mukabagorora, T., Mbonambi, L., Lockhat, Z., Musafiri, A., & Kekana, R. M. (2025). A comprehensive scoping review of existing carotid duplex ultrasound scanning and reporting protocols: identifying gaps and opportunities for standardization of practice in low-income countries. In Journal of Ultrasound (Vol. 28, Number 4, pp. 783–801). Springer Science and Business Media Deutschland GmbH. https://doi.org/10.1007/s40477-025-01064-1
Quinn, L., Tryposkiadis, K., Deeks, J., De Vet, H. C. W., Mallett, S., Mokkink, L. B., Takwoingi, Y., Taylor-Phillips, S., & Sitch, A. (2023). Interobserver variability studies in diagnostic imaging: a methodological systematic review. In British Journal of Radiology (Vol. 96, Number 1148). British Institute of Radiology. https://doi.org/10.1259/bjr.20220972
White, S. J., Phua, Q. S., Lu, L., Yaxley, K. L., McInnes, M. D. F., & To, M. S. (2024). Heterogeneity in Systematic Reviews of Medical Imaging Diagnostic Test Accuracy Studies. JAMA Network Open, 7(2), E240649. https://doi.org/10.1001/jamanetworkopen.2024.0649
Zheng, D., He, X., & Jing, J. (2023). Overview of Artificial Intelligence in Breast Cancer Medical Imaging. In Journal of Clinical Medicine (Vol. 12, Number 2). MDPI. https://doi.org/10.3390/jcm12020419





