DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
The amendments filed 06/03/2026 have been accepted and considered in this office action. Claims 1, 2, 4, 18, and 20 have been amended. Claims 1-20 are pending.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-20 have been considered but are moot in view of new grounds of rejection necessitated by the applicant’s amendments to the claims.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-3, 8-12, 15, and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Kroll (US 20220171003 A1) in view of Amthor et al. (hereinafter Amthor) (US 20230181074 A1) in further view of Takayanagi et al. (hereinafter Takayanagi) (US 20170309275 A1)
Regarding claim 1, Kroll discloses:
A medical image diagnostic system comprising:
an image diagnostic apparatus that acquires a medical image (Kroll, Abstract: "magnetic resonance apparatus");
a first detection device that detects, with at least one of a video or a voice, utterance- related information related to an utterance of a subject during an examination of the subject using the image diagnostic apparatus (Kroll, P[0187], "detected via a speech input unit", the patient's acoustic utterance is detected by a microphone integrated into the MRI apparatus during the MR examination, thus, the speech input unit is the first detector and the detected acoustic signal is voice based utterance related information);
a first display device that displays information in an aspect visible to the subject during the examination using the image diagnostic apparatus (Kroll, P[0171], P[0081]: “the automatic interaction with the patient comprises visual and/or tactile communication.”);
a processor (Kroll, P[0054]); and
a memory that stores a program to be executed by the processor (Kroll, P[0054]),
and to-cause--
in a case in which the subject makes an utterance during the examination using the image diagnostic apparatus (Kroll, P[0171], "receive an acoustic utterance 44 of the patient 15 from a speech input unit 31 and process it", Kroll's disclosed automatic interaction occurs during the magnetic resonance examination), causes the first display device to display, during the examination (Kroll, P[0171], Kroll teaches returning the speech processing result to the MRI patient through a display),
Kroll does not explicitly disclose:
a second detection device that detects, with at least one of a video or a voice, response information of the subject to a question to the subject before the examination of the subject using the image diagnostic apparatus is started;
wherein the processor generates subject feature information related to voice generation of the subject based on the response information detected by the second detection device, [[and]]
recognizes an utterance content of the subject based on the utterance-related information detected by the first detection device and the subject feature information
the recognized utterance content visible to the subject
However, Amthor discloses:
a second detection device that detects, with at least one of a video or a voice, response information of the subject to a question to the subject before the examination of the subject using the image diagnostic apparatus is started (Amthor, P[0018]: “a prediction algorithm takes data about a patient acquired before an exam that can be gathered at a hospital or at home”, P[0025]: “the at least one characteristic of the patient comprises one or more of: age, weight, body-mass index, information on a previous diagnosis, physical condition of the patient, psychological condition of the patient, completed questionnaire information, patient feedback.”, P[0019]: “In an example, the at least one sensor data of the patient was acquired by one or more of: camera, microphone”, P[0021]: “the at least one sensor data of the patient comprises one or more of: …voice data”, Amthor teaches acquiring patient data before the MRI examination, including completed questionnaire information and patient feedback, and teaches acquiring patient sensor data by a camera of microphone, including voice data.);
It would have been prima facie obvious to one of ordinary skill in art before the effective filing date of the claimed invention to use the patient monitoring and intervention techniques as taught by Amthor in the system of Kroll in order to determine when patient-directed intervention is appropriate based on monitored patient condition during the examination.
The combination of Kroll and Amthor does not explicitly disclose:
wherein the processor generates subject feature information related to voice generation of the subject based on the response information detected by the second detection device, [[and]]
recognizes an utterance content of the subject based on the utterance-related information detected by the first detection device and the subject feature information
the recognized utterance content visible to the subject However, Takayanagi discloses:
wherein the processor generates subject feature information related to voice generation of the subject based on the response information detected by the second detection device (Takayanagi, P[0124], generates subject-dependent feature/reference information from that user's prior speech/lip training input), [[and]]
recognizes an utterance content of the subject based on the utterance-related information detected by the first detection device (Takanayagi, P[0116], Takayanagi recognizes the current utterance content from the detected audio feature signal) and the subject feature information (Takanayagi, P[0079], Takayanagi performs recognition by comparing current utterance-derived features against previously stored feature/reference information)
the recognized utterance content visible to the subject (Takanayagi, P[0012]: "render the combined dictation as text on a display", explicitly teaches displaying recognized dictation itself as text). It would have been prima facie obvious to one of ordinary skill in art before the effective filing date of the claimed invention to use the subjected dependent speech recognition techniques as taught by Takayanagi in the system of Kroll as modified by Amthor in order to improve recognition of patient utterances using information specific to the patient.
Regarding claim 2, the combination of Kroll, Amthor, and Takayanagi discloses the medical image diagnostic system according to claim 1.
The combination further discloses:
wherein the first detection device is disposed in the image diagnostic apparatus, and includes at least one of a first camera that captures a video including at least a lip part of a face region of the subject during the examination using the image diagnostic apparatus, or a first microphone that detects a voice uttered by the subject (Kroll, Fig 1, P[0173]: "the speech input unit 31 and/or the output unit 33 are integrated in the local receiving antenna 26.", Fig. 2, P[0175]: "receive an acoustic utterance 44 of the patient 15 via a microphone" (explicitly permits the speech-input unit to be integrated into the MRI receiving antenna and implements that speech-input unit with a microphone receiving the patients acoustic utterance)).
Regarding claim 3, the combination of Kroll, Amthor, and Takayanagi discloses the medical image diagnostic system according to claim 2.
The combination further discloses:
wherein the utterance-related information is a lip movement of the subject acquired from the video captured by the first camera (Abstract, Takayanagi, "video of lip motion", obtains video representing user's lip movement during utterance and the resulting video stream is used for video-based speech recognition).
Regarding claim 8, the combination of Kroll, Amthor, and Takayanagi disclose the medical image diagnostic system according to claim 1.
The combination further discloses:
wherein the processor trains a first machine learning model dedicated to the subject based on the response information detected by the second detection device (Takayanagi, P[0004]: "the service can “learn” the user's unique characteristics", P[0079]: "user speaking a training sequence", "The reference pattern models and/or predetermined feature signals may have been generated based upon a user speaking a training sequence" (Takayanagi generates subject dependent recognition models/features from training information obtained from the particular user, corresponding to training a recognition model dedicated to that subject based on the subject's response/training information)), and inputs the utterance-related information detected by the first detection device to the trained first machine learning model, to acquire the utterance content recognized by the first machine learning model (Takayanagi, P[0079]: "matching the feature signal", Takanayagi feeds features obtained from the newly received audio/video utterance into its recognition process and compares them with the stored trained feature/reference information to generate textual dictation).
Regarding claim 9, the combination of Kroll, Amthor, and Takayanagi disclose the medical image diagnostic system according to claim 8.
The combination further discloses:
wherein the subject feature information includes parameters optimized in a process of training the first machine learning model (Amthor, P[0094]: "Patient feedback on the experienced stress and movement for that patient and other patients can also be provided for training the internal machine learning algorithms") based on response information of the subject (Takayanagi, P[0079]: "the reference pattern models and/or predetermined feature signals may have been generated based upon a user speaking a training sequence. ", Takayanagi generates the subject dependent recognition information from training information provided by that particular user).
Regarding claim 10, the combination of Kroll, Amthor, and Takayanagi disclose the medical image diagnostic system according to claim 8.
The combination further discloses:
wherein a second machine learning model that has been trained through machine learning in advance based on a training data set consisting of utterance-related information related to utterances of a plurality of people is provided (Takanayagi, P[0004]: “However, in subject independent (SI) lip reading services, errors can occur due to large variations within lip shapes, skin textures around the mouth, varying speaking speeds and different accents, which could significantly affect the spatiotemporal appearances of a speaking mouth. A recent SI lip reading algorithm developed by Zhou et al. can reportedly achieve recognition rates as high as 92.8%” (Subject Independent models are trained on a plurality of people to handle large variations within lip shapes and skin textures across a great population)), and
the processor inputs the utterance-related information detected by the second detection device to the second machine learning model, to acquire the utterance content recognized by the second machine learning model in a case in which the utterance content recognized by the first machine learning model is not a meaningful content or a certainty degree of the utterance content is less than a threshold value (Takanayagi, P[0109]-P[0110]: “if no prototype signal has a probability greater than a predetermined standard such as, for example, 90%, the portion of the video stream corresponding to this portion of the audio stream (the previous Y time units) is input to the video based recognition module”, “the decision to proceed to 524 could be decided based upon a combination of if the word is a characteristic word and the probability of the prototype signal” (hierarchical logic is described here where the system checks the “probability” (certainty degree) of a result against a “predetermined standard” (threshold). If the result is not suitable, the system proceeds to another recognition module/model. Using a general (SI) model as a fallback for a low-probability specialized (SD) result is a standard application of the logic here.)).
Regarding claim 11, the combination of Kroll, Amthor, and Takayanagi disclose the medical image diagnostic system according to claim 1.
The combination further discloses:
further comprising: a notification device that notifies an operator in an operation room of the image diagnostic apparatus (Kroll, P[0176]: "in a control room 4", Kroll places an operator side input/output arrangement in the control room used for control/monitoring of the MRI apparatus) of the utterance content, wherein the processor outputs the utterance content to the notification device (Kroll, P[0141]: "detected acoustic utterance of the patient is output approximately in real time to the user of the magnetic resonance apparatus", Kroll outputs the patient's detected utterance to the MRI system user/operator).
Regarding claim 12, the combination of Kroll, Amthor, and Takayanagi disclose the medical image diagnostic system according to claim 11.
The combination further discloses:
wherein the notification device is at least one of a second display device that displays characters indicating the utterance content or a speaker that generates a voice indicating the utterance content (Kroll, P[0148]: "presentation of the keyword on a display", teaches presenting speech processing information derived from the patients utterance visually on a display to the MRI system user/operator).
Regarding claim 15, the combination of Kroll, Amthor, and Takayanagi disclose the medical image diagnostic system according to claim 1.
The combination further discloses:
wherein the vital information measurement device measures one or more of a heart rate, a blood pressure, a respiratory rate, a body temperature, an electrocardiogram, or a blood oxygen saturation concentration of the subject (Amthor, P[0021]: "breathing rate data, heart rate data, voice data, skin resistance data, skin temperature data, skin humidity data, skin motion data, body part movement data, blinking frequency data, EEG data").
Regarding claim 17, the combination of Kroll, Amthor, and Takayanagi disclose the medical image diagnostic system according to claim 1.
The combination further discloses:
wherein the image diagnostic apparatus includes a magnetic resonance imaging apparatus, an X-ray CT apparatus, a PET apparatus, a radiation therapy apparatus, or a particle beam therapy apparatus (Kroll, Abstract, magnetic resonance apparatus).
Regarding claim 18, claim 18 recites an operation method corresponding to the medical image diagnostic system associated with claim 1 and is rejected for the same reasons as above.
Regarding claim 19, claim 19 recites an operation method corresponding to the medical image diagnostic system associated with claim 8 and is rejected for the same reasons as above.
Regarding claim 20, claim 20 recites an information processing system corresponding to the medical image diagnostic system associated with claim 1 and is rejected for the same reasons as above.
Claims 4-5, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Kroll (US 20220171003 A1) in view of Amthor et al. (hereinafter Amthor) (US 20230181074 A1) in further view of Takayanagi et al. (hereinafter Takayanagi) (US 20170309275 A1) and Weiss et al. (hereinafter Weiss) (US 20240324889 A1).
Regarding claim 4, the combination of Kroll, Amthor, and Takayanagi disclose the medical image diagnostic system according to claim 1.
The combination does not explicitly disclose:
wherein the second detection device is farther from the image diagnostic apparatus than the first detection device, and includes at least one of a second camera that captures a video including a face region of the subject, or a second microphone that detects a voice uttered by the subject. However, Weiss discloses:
wherein the second detection device is farther from the image diagnostic apparatus than the first detection device (Weiss, P[0004]: "the camera system is mounted relatively far away from the region to be observed" (Weiss teaches positioning the second patient monitoring detector relatively far from MRI examination reading)), and includes at least one of a second camera that captures a video including a face region of the subject (Weiss, "a camera can observe the patient, e.g. the patient's face," (Weiss teaches second camera alternative capturing the subjects face)), or a second microphone that detects a voice uttered by the subject (See alternative mapped above).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to use the camera and display arrangement as taught by Weiss in the system of Kroll as modified by Amthor and Takayanagi in order to monitor the patient while allowing information to be displayed to the patient during the medical imaging examination.
Regarding claim 5, the combination of Kroll, Amthor, Takayanagi, and Weiss disclose the medical image diagnostic system according to claim 4.
The combination further discloses:
wherein the response information is a lip movement of the subject acquired from the video captured by the second camera (Takayanagi, P[0020]: "video signal representative of lip movement" (Takayanagi treats lip movement obtained by the video input device as the information extracted from the video)).
Regarding claim 16, the combination of Kroll, Amthor, and Takayanagi disclose the medical image diagnostic system according to claim 1.
The combination does not explicitly disclose:
wherein the first display device is a projector that performs projection onto a screen visible to the subject, a head-up display, a head-mounted display, a liquid crystal display, or an organic EL display (Weiss, P[0014] P[0095], projector).
However, Weiss discloses:
wherein the first display device is a projector that performs projection onto a screen visible to the subject, a head-up display, a head-mounted display, a liquid crystal display, or an organic EL display (Weiss, P[0014] P[0095], projector).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to use the camera and display arrangement as taught by Weiss in the system of Kroll as modified by Amthor and Takayanagi in order to monitor the patient while allowing information to be displayed to the patient during the medical imaging examination.
Claims 6-7, and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Kroll (US 20220171003 A1) in view of Amthor et al. (hereinafter Amthor) (US 20230181074 A1) in further view of Takayanagi et al. (hereinafter Takayanagi) (US 20170309275 A1) and Palanisamy et al. (hereinafter Palanisamy) (US 20240008783 A1).
Regarding claim 6, the combination of Kroll, Amthor, and Takayanagi disclose the medical image diagnostic system according to claim 1.
The combination of Kroll, Amthor, and Takayanagi does not explicitly disclose:
However, Palanisamy discloses:
further comprising:
a question apparatus that asks a question to the subject, wherein the question apparatus asks a question of a predetermined format (Palanisamy, P[0053]: "questions may be predefined") with at least one of a voice or characters on a monitor screen (Palanisamy, P[0069]: "visual text audio").
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to use the patient questioning and response techniques as taught by Palanisamy in the system of Kroll as modified by Amthor and Takayanagi in order to obtain patient feedback and improve communication with the patient during the medical imaging examination.
Regarding claim 7, the combination of Kroll, Amthor, Takayanagi, and Palanisamy disclose the medical image diagnostic system according to claim 6.
The combination further discloses:
further comprising:
a question apparatus that asks a question to the subject, wherein the question apparatus asks a question of a predetermined format (Palanisamy, P[0053]: "questions may be predefined") with at least one of a voice or characters on a monitor screen (Palanisamy, P[0069]: "visual text audio").
Regarding claim 14, the combination of Kroll, Amthor, and Takayanagi disclose the medical image diagnostic system according to claim 1.
The combination does not explicitly disclose:
further comprising:
a vital information measurement device that measures vital information of the subject during the examination of the subject using the image diagnostic,
wherein the processor determines whether or not a reply to the utterance content is necessary, based on the utterance content and the measured vital information, and creates a reply sentence corresponding to the utterance content to cause the first display device to display the reply sentence in a case in which it is determined that the reply is.
However, Palanisamy discloses:
further comprising:
a vital information measurement device that measures vital information of the subject during the examination of the subject using the image diagnostic apparatus (Palanisamy, P[0034]: "ECG, EMG, Temperature, BP, and SPO2", teaches monitoring physiological/vital info in its medical imaging patient monitoring system),
wherein the processor determines whether or not a reply to the utterance content is necessary, based on the utterance content (Palanisamy, P[0089], processes patient feedback conveyed as speech commands using speech-to-text, thus, supplying processing of the recognized utterance content) and the measured vital information (Palanisamy, P[0054], "output of the patient condition module is used to select questions from these groups"), and creates a reply sentence corresponding to the utterance content (Palanisamy, P[0073, dialog generator) to cause the first display device to display the reply sentence in a case in which it is determined that the reply is necessary (Palanisamy, P[0069], visual text audio).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date to use the patient questioning and response techniques as taught by Palanisamy in the system of Kroll as modified by Amthor and Takayanagi in order to obtain patient feedback and improve communication with the patient during the medical imaging examination.
Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable Kroll (US 20220171003 A1) in view of Amthor et al. (hereinafter Amthor) (US 20230181074 A1) in further view of Takayanagi et al. (hereinafter Takayanagi) (US 20170309275 A1) and Miller et al. (hereinafter Miller) (US 20060074286 A1).
Regarding claim 13, the combination of Kroll, Amthor, and Takayanagi disclose the medical image diagnostic system according to claim 3.
The combination does not explicitly disclose:
wherein the processor determines whether or not the lip movement of the subject hinders the examination of the subject using the image diagnostic apparatus in a case in which next utterance-related information is not detected for a certain time or longer after the first detection device detects the utterance-related information, and causes the first display device to display characters prompting the subject to make an utterance in a case in which it is determined that the lip movement of the subject does not hinder the examination of the subject using the image diagnostic apparatus However, Miller discloses:
wherein the processor determines whether or not the lip movement of the subject hinders the examination of the subject using the image diagnostic apparatus in a case in which next utterance-related information is not detected for a certain time or longer after the first detection device detects the utterance-related information, and causes the first display device to display characters prompting the subject to make an utterance in a case in which it is determined that the lip movement of the subject does not hinder the examination of the subject using the image diagnostic apparatus (Miller, Abstract: "prompt the patient, using the patient display, to perform a bodily action that facilitates the scan", P[0028]: "In accordance with various embodiments of the invention, the promptings are based on the actions of the patient. The patient's actions are monitored. Based on these monitored actions" (Monitoring patient actions (includes lip movement/speech) and determining if those actions hinder the scan via artifacts is taught. “Timeout” logic is also used here where, if a specific action (or lack thereof) is detected for a certain time the system triggers the next step. The specific hardware (patient display) and the logic of “timing” the prompt so it only occurs when it “facilitates the scan” (ex. When it doesn’t hinder it) is taught. Breath holding mentioned in this reference can be obviously applied to and reads on making an utterance).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the patient monitoring and prompting techniques as taught by Miller in the system of Kroll as modified by Amthor and Takayanagi in order to monitor patient actions and provide prompts that facilitate the medical imaging examination.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHASHIDHAR S MANOHARAN whose telephone number is (571)272-6772. The examiner can normally be reached M-F 8:00-4:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHASHIDHAR SHANKAR MANOHARAN/Examiner, Art Unit 2655
/DOUGLAS GODBOLD/Primary Examiner, Art Unit 2655