DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
displaying, concurrently with the generating of the text transcription, the text transcription 1, 13
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
1. Claims 1-3, 7, 9, 10, 12-14, 18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Bowers et al US 2023/0046658 A1 (“Bowers”) in view of Ma et al US 2025/0173043 A1 (“Ma”)
Per claim 1, Bowers discloses a computer-implemented method comprising:
causing generating, using a trained model, a plurality of candidate responses to a first portion of audio content (fig. 4B; The speech recognition engine(s) 120A1 and/or 120A2 can process, using speech recognition model(s) 120A, the audio data that captures the spoken input to generate recognized text…. the speech recognition model(s) 120A include a single speech recognition model that is trained to process audio data …, para. [0038]; As described herein, in some implementations, suggestions(s) that are responsive to spoken input of additional participant(s) detected at the client device 410 may only be rendered responsive to determining a given user of the client device 410 is an intended target of the spoken input. In some implementations, the client device 410 can generate one or more suggestions 456B1-456B4 …, para. [0096]);
receiving a selection of a candidate response in the plurality of candidate responses (Upon receiving further user interface input at the client device 410 directed to one of suggestions 456B1-456B4 …, para. [0096]); and
causing generating, using a trained text-to-speech model, a second portion of audio content, wherein the second portion of audio content comprises a spoken version of the selected candidate response (para. [0009]; para. [0028]; Upon receiving further user interface input at the client device 410 directed to one of suggestions 456C1-456C4, the client device 410 can generate synthesized speech, using Tim's speech embedding established in FIG. 4A, that includes the candidate textual segment of the selected suggestion …, para. [0096])
Bowers does not explicitly disclose causing generating, using a trained large language model (LLM), a plurality of candidate responses
However, this feature is taught by Ma (fig. 3; fig. 4A; para. [0004]; para. [0013]; an input prompt/query for an LLM is received from a user device…. For example, if the query is a voice query the system can perform automatic speech recognition (ASR) to convert the voice query into textual format., para. [0094]-[0097]; At block 606, the system causes, based on the first LLM output, the set of UI elements to be rendered at a user device. For instance, the system can generate the set of UI elements based on the first LLM output …, para. [0098])
It would have been obvious to one of ordinary skill in the art before the effective filing of the invention to combine the teachings of Ma with the method of Bowers in arriving at the missing features of Bowers, because such combination would have resulted in providing flexibility and intuitiveness, while guiding a user through a structured and easy to use process with a user interface (Ma, para. [0040])
Per claim 2, Bowers in view of Ma discloses the computer-implemented method of claim 1,
Bowers discloses further comprising: causing generating, using a trained speech-to-text model, a text transcription of the first portion of audio content (para. [0038]); and
displaying, concurrently with the generating of the text transcription, the text transcription (fig. 4B; para. [0085]; para. [0096]).
Per claim 3, Bowers in view of Ma discloses the computer-implemented method of claim 2,
Ma discloses wherein the text transcription is displayed on a display of a mixed reality device (para. [0054]-[0055]).
Per claim 7, Bowers in view of Ma discloses the computer-implemented method of claim 3,
Bowers discloses wherein the second portion of audio content is generated on a device other than the mixed reality device (fig. 1; para. [0037]).
Ma discloses wherein the second portion of audio content is generated on a device other than the mixed reality device (fig. 1; para. [0054]-[0055]; para. [0060]).
Per claim 9, Bowers in view of Ma discloses the computer-implemented method of claim 1,
Bowers discloses wherein the plurality of candidate responses is each generated in text form (fig. 4B).
Per claim 10, Bowers in view of Ma discloses the computer-implemented method of claim 1,
Ma discloses wherein the plurality of candidate responses is each generated in audio form (para. [0055]).
Per claim 12, Bowers discloses a non-transitory computer-readable medium storing a program, which when executed by a computer, configures the computer to:
cause generating, using a trained model, a plurality of candidate responses to a first portion of audio content (fig. 4B; The speech recognition engine(s) 120A1 and/or 120A2 can process, using speech recognition model(s) 120A, the audio data that captures the spoken input to generate recognized text…. the speech recognition model(s) 120A include a single speech recognition model that is trained to process audio data …, para. [0038]; As described herein, in some implementations, suggestions(s) that are responsive to spoken input of additional participant(s) detected at the client device 410 may only be rendered responsive to determining a given user of the client device 410 is an intended target of the spoken input. In some implementations, the client device 410 can generate one or more suggestions 456B1-456B4 …, para. [0096]);
receive a selection of a candidate response in the plurality of candidate responses (Upon receiving further user interface input at the client device 410 directed to one of suggestions 456B1-456B4 …, para. [0096]); and
cause generating, using a trained text-to-speech model, a second portion of audio content, wherein the second portion of audio content comprises a spoken version of the selected candidate response (para. [0009]; para. [0028]; Upon receiving further user interface input at the client device 410 directed to one of suggestions 456C1-456C4, the client device 410 can generate synthesized speech, using Tim's speech embedding established in FIG. 4A, that includes the candidate textual segment of the selected suggestion …, para. [0096])
Bowers does not explicitly disclose cause generating, using a trained large language model (LLM), a plurality of candidate responses
However, this feature is taught by Ma (fig. 3; fig. 4A; para. [0004]; para. [0013]; an input prompt/query for an LLM is received from a user device…. For example, if the query is a voice query the system can perform automatic speech recognition (ASR) to convert the voice query into textual format., para. [0094]-[0097]; At block 606, the system causes, based on the first LLM output, the set of UI elements to be rendered at a user device. For instance, the system can generate the set of UI elements based on the first LLM output …, para. [0098])
It would have been obvious to one of ordinary skill in the art before the effective filing of the invention to combine the teachings of Ma with the medium of Bowers in arriving at the missing features of Bowers, because such combination would have resulted in providing flexibility and intuitiveness, while guiding a user through a structured and easy to use process with a user interface (Ma, para. [0040])
Per claim 13, Bowers in view of Ma discloses the non-transitory computer-readable medium of claim 12,
Bowers discloses wherein the program, when executed by the computer, further configures the computer to: cause generating, using a trained speech-to-text model, a text transcription of the first portion of audio content (para. [0038]); and
display, concurrently with the generating of the text transcription, the text transcription (fig. 4B; para. [0085[; para. [0096]).
Per claim 14, Bowers in view of Ma discloses the non-transitory computer-readable medium of claim 13,
Ma discloses wherein the text transcription is displayed on a display of a mixed reality device (para. [0054]-[0055]).
Per claim 18, Bowers in view of Ma discloses the non-transitory computer-readable medium of claim 14,
Bowers discloses wherein the second portion of audio content is generated on a device other than the mixed reality device (fig. 1; para. [0037]).
Ma discloses wherein the second portion of audio content is generated on a device other than the mixed reality device (fig. 1; para. [0054]-[0055]; para. [0060]).
Per claim 20, Bowers discloses a system comprising:
a processor (para. [0111]); and
a non-transitory computer-readable medium storing a set of instructions, which when executed by the processor, configure the system to: cause generating, using a trained model, a plurality of candidate responses to a first portion of audio content (fig. 4B; The speech recognition engine(s) 120A1 and/or 120A2 can process, using speech recognition model(s) 120A, the audio data that captures the spoken input to generate recognized text…. the speech recognition model(s) 120A include a single speech recognition model that is trained to process audio data …, para. [0038]; As described herein, in some implementations, suggestions(s) that are responsive to spoken input of additional participant(s) detected at the client device 410 may only be rendered responsive to determining a given user of the client device 410 is an intended target of the spoken input. In some implementations, the client device 410 can generate one or more suggestions 456B1-456B4 …, para. [0096]; para. [0111]);
receive a selection of a candidate response in the plurality of candidate responses (Upon receiving further user interface input at the client device 410 directed to one of suggestions 456B1-456B4 …, para. [0096]); and
cause generating, using a trained text-to-speech model, a second portion of audio content, wherein the second portion of audio content comprises a spoken version of the selected candidate response (para. [0009]; para. [0028]; Upon receiving further user interface input at the client device 410 directed to one of suggestions 456C1-456C4, the client device 410 can generate synthesized speech, using Tim's speech embedding established in FIG. 4A, that includes the candidate textual segment of the selected suggestion …, para. [0096])
Bowers does not explicitly disclose causing generating, using a trained large language model (LLM), a plurality of candidate responses
However, this feature is taught by Ma (fig. 3; fig. 4A; para. [0004]; para. [0013]; an input prompt/query for an LLM is received from a user device…. For example, if the query is a voice query the system can perform automatic speech recognition (ASR) to convert the voice query into textual format., para. [0094]-[0097]; At block 606, the system causes, based on the first LLM output, the set of UI elements to be rendered at a user device. For instance, the system can generate the set of UI elements based on the first LLM output …, para. [0098])
It would have been obvious to one of ordinary skill in the art before the effective filing of the invention to combine the teachings of Ma with the system of Bowers in arriving at the missing features of Bowers, because such combination would have resulted in providing flexibility and intuitiveness, while guiding a user through a structured and easy to use process with a user interface (Ma, para. [0040]).
2. Claims 4, 5, 15 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Bowers in view of Ma as applied to claims 3 and 14 above, and further in view of Bekker et al US 2023/0197064 A1 (“Bekker”)
Per claim 4, Bowers in view of Ma discloses the computer-implemented method of claim 3,
Bowers in view of Ma does not explicitly disclose wherein the trained speech-to-text model executes on the mixed reality device
However, this feature is taught by Bekker (fig. 2; para. [0025]; para. [0060])
It would have been obvious to one of ordinary skill in the art before the effective filing of the invention to combine the teachings of Bekker with the method of Bowers in view of Ma in arriving at the missing features of Bowers in view of Ma, because such combination would have resulted in significantly improving the overall ability for a social network system to perform automated computer tasks using speech recognition (Bekker, para. [0002]; para. [0018])
Per claim 5, Bowers in view of Ma discloses the computer-implemented method of claim 3,
Bowers in view of Ma does not explicitly disclose wherein the trained speech-to-text model executes on a device other than the mixed reality device.
However, this feature is taught by Bekker (fig. 2; para. [0025]; para. [0060])
It would have been obvious to one of ordinary skill in the art before the effective filing of the invention to combine the teachings of Bekker with the method of Bowers in view of Ma in arriving at the missing features of Bowers in view of Ma, because such combination would have resulted in significantly improving the overall ability for a social network system to perform automated computer tasks using speech recognition (Bekker, para. [0002]; para. [0018]).
Per claim 15, Bowers in view of Ma discloses the non-transitory computer-readable medium of claim 14,
Bowers in view of Ma does not explicitly disclose wherein the trained speech-to-text model executes on the mixed reality device
However, this feature is taught by Bekker (fig. 2; para. [0025]; para. [0060])
It would have been obvious to one of ordinary skill in the art before the effective filing of the invention to combine the teachings of Bekker with the medium of Bowers in view of Ma in arriving at the missing features of Bowers in view of Ma, because such combination would have resulted in significantly improving the overall ability for a social network system to perform automated computer tasks using speech recognition (Bekker, para. [0002]; para. [0018])
Per claim 16, Bowers in view of Ma discloses the non-transitory computer-readable medium of claim 14,
Bowers in view of Ma does not explicitly disclose wherein the trained speech-to-text model executes on a device other than the mixed reality device.
However, this feature is taught by Bekker (fig. 2; para. [0025]; para. [0060])
It would have been obvious to one of ordinary skill in the art before the effective filing of the invention to combine the teachings of Bekker with the medium of Bowers in view of Ma in arriving at the missing features of Bowers in view of Ma, because such combination would have resulted in significantly improving the overall ability for a social network system to perform automated computer tasks using speech recognition (Bekker, para. [0002]; para. [0018])
3. Claims 6 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Bowers in view of Ma as applied to claims 3 and 14 above, and further in view of Powderly et al US 20220171453 A1 (“Powderly”)
Per claim 6, Bowers in view of Ma discloses the computer-implemented method of claim 3,
Bowers in view of Ma does not explicitly disclose wherein the selection of the candidate response is performed using an eye movement tracker function of the mixed reality device
However, this feature is taught by Powderly (para. [0056]-[0057])
It would have been obvious to one of ordinary skill in the art before the effective filing of the invention to combine the teachings of Powderly with the method of Bowers in view of Ma in arriving at the missing features of Bowers in view of Ma, because such combination would have resulted in providing significant improvement in speed and convenience when a user is performing tasks (Powderly, para. [0280]).
Per claim 17, Bowers in view of Ma discloses the non-transitory computer-readable medium of claim 14,
Bowers in view of Ma does not explicitly disclose wherein the selection of the candidate response is performed using an eye movement tracker function of the mixed reality device
However, this feature is taught by Powderly (para. [0056]-[0057])
It would have been obvious to one of ordinary skill in the art before the effective filing of the invention to combine the teachings of Powderly with the medium of Bowers in view of Ma in arriving at the missing features of Bowers in view of Ma, because such combination would have resulted in providing significant improvement in speed and convenience when a user is performing tasks (Powderly, para. [0280]).
4. Claims 8 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Bowers in view of Ma as applied to claims 3 and 14 above, and further in view of Waibel et al US 2023/0186899 A1 (“Waibel”)
Per claim 8, Bowers in view of Ma discloses the computer-implemented method of claim 1,
Bowers discloses causing generating, using a second trained speech-to-text model, a text transcription of the first portion of audio content (para. [0038]),
Bowers in view of Ma does not explicitly disclose causing generating, using a second trained speech-to-text model, a translated text transcription of the first portion of audio content, wherein the second trained speech-to-text model is trained to generate text in a first human language from input audio content in a second human language; or displaying, concurrently with the generating of the translated text transcription, the translated text transcription
However, these features are suggested by Waibel:
causing generating, using a second trained speech-to-text model, a translated text transcription of the first portion of audio content (para. [0003]; para. [0032]; para. [0062])
wherein the second trained speech-to-text model is trained to generate text in a first human language from input audio content in a second human language (para. [0003]; para. [0032]; para. [0062]); and
displaying, concurrently with the generating of the translated text transcription, the translated text transcription (para. [0003]; para. [0032]; para. [0062])
It would have been obvious to one of ordinary skill in the art before the effective filing of the invention to try to combine the teachings of Waibel with the method of Bowers in view of Ma in arriving at the missing features of Bowers in view of Ma, because such combination would have resulted in improving recognition for unusual vocabularies as well as typical accents and noise conditions found in different deployments or venues (Waibel, para. [0032]).
Per claim 19, Bowers in view of Ma discloses the non-transitory computer-readable medium of claim 12,
Bowers discloses wherein the program, when executed by the computer, further configures the computer to: cause generating, using a second trained speech-to-text model, a text transcription of the first portion of audio content (para. [0038]),
Bowers in view of Ma does not explicitly disclose causing generating, using a second trained speech-to-text model, a translated text transcription of the first portion of audio content, wherein the second trained speech-to-text model is trained to generate text in a first human language from input audio content in a second human language or display, concurrently with the generating of the translated text transcription, the translated text transcription
However, these features are suggested by Waibel:
when executed by the computer, further configures the computer to: cause generating, using a second trained speech-to-text model, a text transcription of the first portion of audio content (para. [0003]; para. [0032]; para. [0062])
wherein the second trained speech-to-text model is trained to generate text in a first human language from input audio content in a second human language (para. [0003]; para. [0032]; para. [0062]); and
display, concurrently with the generating of the translated text transcription, the translated text transcription (para. [0003]; para. [0032]; para. [0062])
It would have been obvious to one of ordinary skill in the art before the effective filing of the invention to try to combine the teachings of Waibel with the medium of Bowers in view of Ma in arriving at the missing features of Bowers in view of Ma, because such combination would have resulted in improving recognition for unusual vocabularies as well as typical accents and noise conditions found in different deployments or venues (Waibel, para. [0032])
5. Claim 11 is ejected under 35 U.S.C. 103 as being unpatentable over Bowers in view of Ma as applied to claim 1 above, and further in view of Volkan et al US 2024/0087199 A1 (“Volkan”) and Wang et al US 2023/0237723 A1 (“Wang”)
Per claim 11, Bowers in view of Ma discloses the computer-implemented method of claim 1,
Bowers in view of Ma does not explicitly disclose causing generating, using a trained avatar generation model, a portion of video content, wherein the portion of video content comprises an avatar portrayed as speaking the selected candidate response
However, this feature is suggested by Volkan that discloses causing generating, using a avatar generation model, a portion of video content, wherein the portion of video content comprises an avatar portrayed as speaking the selected candidate response (para. [0003]; para. [0017]-[0019]; para. [0021]; claim 8)
Bowers in view of Ma does not explicitly disclose the use of a trained avatar generation model
However, this feature is taught by Wang (para. [0007]; para. [0057])
It would have been obvious to one of ordinary skill in the art before the effective filing of the invention to try to combine the teachings of Volkan with the method of Bowers in view of Ma in arriving at the missing features of Bowers in view of Ma, as well as to combine the teachings of Wang with the method of Bowers in view of Ma and Volkan in arriving at the missing features of Bowers in view of Ma and Volkan, because such combination would have resulted in helping a user better understand how to pronounce selected text (Volkan, para. [0019]), as well as in enhancing use convenience for users and improving user experience (Wang, para. [0055]-[0057])
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See PTO 892 form.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to OLUJIMI A ADESANYA whose telephone number is (571)270-3307. The examiner can normally be reached Monday-Friday 8:30-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at 571-272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/OLUJIMI A ADESANYA/Primary Examiner, Art Unit 2658