Method and system for the computer-aided processing of medical images
Abstract
Methods and systems described enable automatic generation of a significant portion of or all of a clinical report (e.g., radiology report), using multimodal models trained on image and language data. Methods described can transform unstructured language and image information into findings, as well as an accurate and comprehensive clinical report, in a designated style (e.g., writing style). The methods and systems described thus significantly improve performance in generation and processing of clinical reports, in relation to time saved per clinical shift, dictation effort, medical billing, and other performance factors.
Claims
exact text as granted — not AI-modifiedWe claim:
1 . A method comprising:
generating a trained model upon training an artificial intelligence model, comprising a language model component, a vision model component, and an adapter between the language model component and the vision model component; at a computing system comprising an interface with a Picture Archiving and Communication System (PACS), receiving a request from a radiologist to generate a report associated with a session with a patient; and returning the report with a level of completion above a threshold level of completion, wherein the report is returned in a writing style of the radiologist, and wherein returning the report comprises:
transforming a set of images generated during the session and retrieved from the PACS into a set of image representations,
returning a set of radiology outputs upon processing the set of image representations and text data generated from the session with the trained model,
integrating the set of radiology outputs into a draft of the report; and
transmitting the draft of the report to the radiologist.
2 . The method of claim 1 , wherein the threshold level of completion is 75%.
3 . The method of claim 1 , wherein the set of radiology outputs comprises a set of annotations for the set of images.
4 . The method of claim 1 , wherein training the artificial intelligence model comprises a first stage involving aligning the language model component and the vision model component using a training dataset comprising image data paired with clinical reports.
5 . The method of claim 4 , wherein training the artificial intelligence model comprises a second stage of training comprising passing of information the adapter between the language model component and the vision model component, while fixing the language model component and the vision model component.
6 . The method of claim 1 , wherein the set of images comprises positron emission tomography (PET)/computed tomography (CT) images and mammography images.
7 . The method of claim 1 , wherein integrating the set of radiology outputs into the draft of the report comprises integrating a set of measurements returned from the trained model into the draft of the report.
8 . The method of claim 7 , wherein the set of measurements comprises a tumor size measurement.
9 . The method of claim 1 , wherein integrating the set of radiology outputs into the draft of the report comprises integrating an image of the set of images into the draft of the report.
10 . The method of claim 1 , further comprising:
receiving an input indicative of an error in the draft of the report, from the radiologist, at the computing system; and returning an indication at a reporting platform of the computing system that the error has been corrected in an updated draft of the report.
11 . The method of claim 1 , further comprising: detecting an anomaly associated with a clinical indication upon processing the set of images with the trained model, retrieving a set of candidate actions to perform based upon the clinical indication, and executing an action of the set of candidate actions, wherein the action comprises administering care according to a critical results workflow corresponding to the clinical indication.
12 . A system comprising:
a reporting platform comprising a speech recognition system and a user interface; an interface with a Picture Archiving and Communication System (PACS); and a computing system storing:
a trained model comprising a language model component, a vision model component, and an adapter between the language model component and the vision model component, and
computer-readable instructions in non-transitory computer-readable media, that when executed perform:
receiving a request from a radiologist to generate a report associated with a session with a patient,
transforming a set of images generated during the session and retrieved from the PACS into a set of image representations,
returning a set of radiology outputs upon processing the set of image representations and text data generated from the session with the trained model,
integrating the set of radiology outputs into a draft of the report, and
transmitting the draft of the report to the radiologist, wherein the draft of the report comprises a level of completion above a threshold level of completion, and wherein the draft of the report is returned in a writing style of the radiologist.
13 . The system of claim 12 , wherein the set of radiology outputs comprises a set of annotations for the set of images.
14 . The system of claim 12 , wherein the adapter is structured to pass information between the language model component and the vision model component with an attention mechanism.
15 . The system of claim 12 , wherein the trained model comprises a decoder-only model.
16 . The system of claim 12 , wherein the set of radiology outputs comprises a set of findings.
17 . The system of claim 12 , wherein the writing style comprises a stylistic element comprising a word choice of the radiologist.
18 . The system of claim 12 , wherein the language component comprises a large language model (LLM).
19 . The system of claim 12 , wherein the set of images comprises positron emission tomography (PET)/computed tomography (CT) images.
20 . The system of claim 12 , wherein the threshold level of completion is 85%.Join the waitlist — get patent alerts
Track US2026011422A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.