Prosecution Insights
Last updated: August 18, 2026
Application No. 18/459,536

COMPUTATIONALLY CUSTOMIZING INSTRUCTIONAL CONTENT

Final Rejection §103§112
Filed
Sep 01, 2023
Priority
May 20, 2021 — continuation of 11/771,977
Examiner
CHEN, KUANG FU
Art Unit
2143
Tech Center
2100 — Computer Architecture & Software
Assignee
Microsoft Technology Licensing, LLC
OA Round
2 (Final)
80%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 80% — above average
80%
Career Allowance Rate
216 granted / 270 resolved
+25.0% vs TC avg
Strong +68% interview lift
Without
With
+68.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
26 currently pending
Career history
295
Total Applications
across all art units

Statute-Specific Performance

§101
16.9%
-23.1% vs TC avg
§103
50.3%
+10.3% vs TC avg
§102
11.1%
-28.9% vs TC avg
§112
15.2%
-24.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 270 resolved cases

Office Action

§103 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The Amendment filed 5/5/2026 has been entered. Claims 1, 12, 15, 17, and 20 were amended. Claims 1-20 are pending. Claim Objections Claim 13 is objected to because of the following informalities: The phrase the audiovisual data of the instruction should read the audiovisual data of the instructor, consistent with the recitation in claim 12 of computer-generated audiovisual data of the instructor and with the remainder of claim 13 (images of the instructor and a voice of the instructor). Appropriate correction is required. Claim Rejections - 35 USC 112(a) The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 12-19 are rejected under 35 U.S.C. 112(a) as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, at the time the application was filed, had possession of the claimed invention. Independent claim 12, as amended, recites in its concluding wherein clause that the computer-generated audiovisual data is generated via a multimodal model based on audiovisual features of the instructor. Under the broadest reasonable interpretation (MPEP 2111), a multimodal model is a single computational model that operates jointly across more than one data modality, here the audio and the visual modalities, and audiovisual features are a joint or combined feature representation spanning both of those modalities. The specification does not reasonably convey possession of this subject matter. Neither the term "multimodal" nor the term "audiovisual features" appears anywhere in the specification. The specification nowhere describes a single model that jointly processes both audio and video, and nowhere describes deriving, extracting, or operating upon any combined audiovisual feature representation of the instructor. What the specification instead describes is the opposite architecture: two separate, single-modality models. The synthetic media application 142 is disclosed as including a separate audio model 144, which generates audible output that mimics the voice of the instructor, and a distinct video model 146, which generates video output that mimics the appearance of the instructor (specification paragraph [0035]). The two models are trained separately upon the instructor audiovisual data 154 (specification paragraphs [0042] and [0044]), and are invoked separately in operation, the audio model receiving input and outputting words in the voice of the instructor and the video model separately generating images of the instructor uttering those words, after which the separately generated audio and video are synced (specification paragraphs [0035] and [0052]). The claim-support paragraph [0081] confirms that the disclosed computer-implemented model "includes an audio model ... and a video model." Two separate single-modality models that are trained and executed independently and then synced are not a single multimodal model operating on a joint audiovisual feature representation. To the extent the summary paragraphs describe a computer-implemented model that outputs both audio content and video content (specification paragraphs [0007] and [0025]), that description does not cure the deficiency, because those paragraphs use "the model" as a collective label for the separately disclosed audio model and video model, as the detailed description makes explicit (specification paragraphs [0035] and [0081]). A collective reference to two single-modality models does not reasonably convey possession of a single multimodal model, nor of the recited audiovisual features on which such a model is claimed to operate. Claim 12 thus defines the audiovisual-generation operation by the result to be achieved, generating computer-generated audiovisual data of the instructor, coupled with a model architecture, "multimodal," that the inventor did not describe possessing. A genus claimed by the function or result it achieves, without a corresponding description of the structure that achieves it, does not satisfy the written description requirement (Ariad Pharmaceuticals, Inc. v. Eli Lilly and Co., 598 F.3d 1336, 1349-51 (Fed. Cir. 2010) (en banc); Abbvie Deutschland GmbH and Co. KG v. Janssen Biotech, Inc., 759 F.3d 1285, 1300-01 (Fed. Cir. 2014)). Because the limitation "a multimodal model based on audiovisual features of the instructor" was introduced into claim 12 by the amendment filed May 5, 2026, and is not reasonably conveyed by the original disclosure, it constitutes new matter (In re Rasmussen, 650 F.2d 1212, 1214 (CCPA 1981); MPEP 2163.06). The new matter is confined to the claims rather than to the specification, so it is addressed by this written description rejection under 35 U.S.C. 112(a), and no separate objection to the specification under 35 U.S.C. 132 is made. Applicant is required either to point out with specificity where the original disclosure describes a single multimodal model operating on audiovisual features of the instructor, or to cancel the limitation. For the purposes of examination, claim 12 limitations wherein the computer-generated audiovisual data is generated via a multimodal model based on audiovisual features of the instructor is interpreted as herein the computer-generated audiovisual data is generated via an audio model and a video model based on an audiovisual data of the instructor. Dependent claims 13-19 depend, directly or indirectly, from claim 12 and incorporate the same "multimodal model based on audiovisual features" limitation, and none adds disclosure curing the deficiency. To the contrary, claim 14 recites that the computer-generated images of the instructor are generated by a computer-implemented model that has been trained based upon video of the instructor, which reflects the separately disclosed single-modality video model rather than a single multimodal model. Claims 13 through 19 are therefore rejected under 35 U.S.C. 112(a) for lack of written description for the same reasons as claim 12. Claims 12-19 are rejected under 35 U.S.C. 112(a) as failing to comply with the enablement requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to enable one skilled in the art to which it pertains, or with which it is most nearly connected, to make and/or use the invention. Whether a disclosure would require undue experimentation is determined by weighing the factors set forth in In re Wands, 858 F.2d 731, 737 (Fed. Cir. 1988). Those factors have been considered, and each is addressed below with respect to the full scope of the claimed generation of computer-generated audiovisual data of the instructor via a multimodal model based on audiovisual features of the instructor. (a) The breadth of the claims. Claim 12 recites generating the computer-generated audiovisual data via a multimodal model based on audiovisual features of the instructor, without limitation as to architecture. The limitation reads on a genus encompassing any multimodal-model architecture operating on any joint audiovisual feature representation of the instructor. A claim of this breadth requires a correspondingly broad enabling disclosure. (b) The nature of the invention. The invention is a computer-implemented method that synthesizes, or deepfakes, audiovisual media depicting a specific human instructor, customized based upon user data pertaining to performance of an activity. (c) The state of the prior art. As of the effective filing date of May 20, 2021, generation of synthetic video of a person via generative neural networks and synthesis of a person's voice were established techniques in the art. The record does not, however, establish that a single multimodal model operating on a joint audiovisual feature representation of a particular instructor was a routine, off-the-shelf technique, and in any event the specification supplies no such technique. (d) The level of one of ordinary skill. The level of ordinary skill is relatively high, corresponding to an advanced degree, or equivalent experience, in machine learning, computer vision, or media synthesis. A high level of skill reduces, but does not eliminate, the disclosure required, and cannot supply subject matter that the specification omits entirely. (e) The level of predictability in the art. Machine learning and software are generally predictable arts. Predictability, however, operates upon a disclosed starting point, permitting a person of ordinary skill to foresee the results of modifying what has already been taught. Here the specification teaches no multimodal-model embodiment to modify, so predictability does not bridge the gap. (f) The amount of direction or guidance provided by the inventor. The specification provides no direction or guidance for a multimodal model based on audiovisual features. It never mentions a multimodal model or audiovisual features, and discloses no architecture, algorithm, feature definition, training objective, or parameter for one. The only model guidance in the specification is directed to two separate single-modality models, the audio model 144 and the video model 146 (specification paragraphs [0035], [0042], [0044], and [0052]). (g) The existence of working examples. The specification contains no working example, actual or prophetic, of a multimodal model based on audiovisual features. The only worked description of media generation, the modification of a template video reciting "Good job [blank]!" to insert the name of the user, is carried out using the separate audio model and video model (specification paragraph [0035]), not a multimodal model. (h) The quantity of experimentation necessary. To make and use the full scope of the claimed subject matter, a person of ordinary skill would be required to independently conceive and construct a multimodal-model architecture and a joint audiovisual feature representation without any direction from the specification. That undertaking is original development, not the routine implementation of a disclosed teaching, and constitutes undue experimentation. Weighing the Wands factors as a whole, and giving particular weight to the complete absence of direction or guidance (factor (f)), the absence of any working example (factor (g)), and the resulting quantity of experimentation (factor (h)) for the specifically claimed multimodal-model and audiovisual-features scope, the specification does not enable a person of ordinary skill in the art to make and use the full scope of claims 12-19 without undue experimentation. No single factor is dispositive. A best mode rejection is not raised. Because written description and enablement are separate and independent requirements, this enablement rejection is asserted in addition to, and independently of, the written description rejection set forth above. Claim Rejections - 35 USC 112(b) The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-11 and 20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Independent claim 1 recites, in the body, that the user-customized media includes at least one of computer-generated video data that includes an image of the instructor; or computer-generated audio data that includes a voice of the instructor, which under the broadest reasonable interpretation is satisfied by user-customized media that includes only the computer-generated video data, only the computer-generated audio data, or both. The claim then recites wherein the computer-generated video data is generated via a video model based on image features of the instructor, the video model being a first generative neural network; and the computer-generated audio data is generated via an audio model based on voice features of the instructor, the audio model being a second generative neural network. This latter clause conjunctively (joined by "and") recites generation of both the computer-generated video data and the computer-generated audio data, and refers to each data type with the definite article "the." As a result, it cannot be determined with reasonable certainty whether the claim requires both the computer-generated video data and the computer-generated audio data to be present (as the conjunctive "wherein" clause indicates) or only at least one of them (as the "at least one of ... or" body permits). Furthermore, for any embodiment in which only one of the two data types is present, which the body expressly permits, the "wherein" clause's reference to the other, absent, data type ("the computer-generated video data" or "the computer-generated audio data") lacks antecedent basis in that embodiment, so it is unclear whether that portion of the "wherein" clause is a required limitation of the claim. A person having ordinary skill in the art is therefore not apprised of the metes and bounds of the claim with reasonable certainty. See Nautilus, Inc. v. Biosig Instruments, Inc., 572 U.S. 898, 901, 910 (2014); In re Packard, 751 F.3d 1307 (Fed. Cir. 2014); MPEP 2173.05(e) and 2173.02. For clarity, this rejection does not rest on the use of the alternative connective in the "at least one of ... or" phrase itself, which is a definite alternative limitation (see SuperGuide Corp. v. DirecTV Enterprises, Inc., 358 F.3d 870, 886-87 (Fed. Cir. 2004)); the indefiniteness arises from the interaction between the conjunctive "wherein" clause and the disjunctive body. For the purposes of examination the said limitations of claim 1 are interpreted as wherein, when the user-customized media includes the computer-generated video data, the computer-generated video data is generated via a video model based on image features of the instructor, the video model being a first generative neural network; and when the user-customized media includes the computer-generated audio data, the computer-generated audio data is generated via an audio model based on voice features of the instructor, the audio model being a second generative neural network. This interpretation is consistent with the specification, which describes an audio model 144 and a video model 146 that are neural networks such as a generative neural network (for example, an autoencoder or a generative adversarial network) and that generate, respectively, audible output mimicking the instructor's voice and video output mimicking the instructor's appearance. See specification paragraph [0035]; see also [0042] and [0044]. Independent claim 20 (a computer-readable storage medium) recites the identical "at least one of ... or" body limitation and the identical conjunctive "wherein" clause reciting a first generative neural network and a second generative neural network, and is indefinite for the same reason and under the same broadest reasonable interpretation set forth above. Dependent claims 2-11 depend, directly or indirectly, from claim 1. They incorporate the indefinite limitations of claim 1 and do not cure the indefiniteness identified above, and are therefore rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, for the same reason. Claim Rejections - 35 USC 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-20 are rejected under 35 U.S.C. 103 over Song et al. (hereinafter Song), US 2014/0120994 A1, in view of Foley et al. (hereinafter Foley), US 2018/0318647 A1, and further in view of Savchenkov et al. (hereinafter Savchenkov), US 2020/0234690 A1. Foley was disclosed in an IDS dated 9/1/2023. Regarding independent claim 1, Song teaches a computing system, comprising: a processor; and memory storing instructions that, when executed by the processor, cause the processor to perform acts comprising (Song: [0082], "a sensing device 10, a storage unit 20, an image processing unit 30, an image output device 40, a club recognition device 50, and a controller M"; [0086], "The database 500 is preferably configured to store identification information of a plurality of virtual golf simulation apparatuses connected to the server S"; the controller M and the server processor 600, together with the storage unit 20 and database 500 that hold the programmed simulation, shot-analysis, and lesson-provision functions, are a processor and memory storing instructions of a computing system): streaming instructional media to a client device for presentation at the client device, where the instructional media includes video of a human instructor setting forth audible instructions with respect to an activity being performed by a user of the client device (Song: [0215], "transmit the extracted information to the virtual golf simulation apparatus"; [0239], "it is possible to provide lesson content while displaying image and voice information regarding a real lesson pro or pro golfer"; the golf lesson content transmitted over the network for presentation at the user's virtual golf simulation apparatus (a client device), presenting the image and voice of a real lesson pro or pro golfer (a human instructor) giving golf instruction (audible instructions) with respect to the golf shots taken by the user (an activity being performed by a user), is streaming instructional media that includes video of a human instructor setting forth audible instructions); generating user-customized media based upon the user data, where the user-customized media includes at least one of computer-generated video data that includes an image of the instructor; or computer-generated audio data that includes a voice of the instructor (Song: [0217], "the customized lesson provision means 300 of the virtual golf simulation apparatus generates customized lesson content to be provided to the user"; [0210], "the database 500 may store information regarding an image and a voice of a virtual lesson pro"; [0215], "extract information regarding lesson content based on the analysis result"; the customized lesson content (user-customized media) generated by the customized lesson provision means based on the analysis result of the user's golf shots (the user data), the content including the image and the voice of the lesson pro (an image of the instructor and a voice of the instructor), is generating user-customized media that includes at least one of video data that includes an image of the instructor or audio data that includes a voice of the instructor; the customized lesson content requires only one of the image or the voice, and Song provides both); and streaming the user-customized media as part of the instructional media to the client device for presentation at the client device (Song: [0215], "transmit the extracted information to the virtual golf simulation apparatus"; [0217], "generates customized lesson content to be provided to the user"; [0236], "customized lesson content is output to the screen by the customized lesson provision means"; transmitting the generated customized lesson content to, and outputting it on the screen of, the user's virtual golf simulation apparatus, as part of the golf lesson content presented to the user, is streaming the user-customized media as part of the instructional media to the client device for presentation). Song does not expressly teach as the instructional media is being streamed to the client device, obtaining user data that pertains to performance of the activity by the user. However, Foley teaches as the instructional media is being streamed to the client device, obtaining user data that pertains to performance of the activity by the user (Foley: [0008], "displaying information about available live and archived cycling classes that can be accessed by a first user using a first stationary bike via a digital communication network on a display screen at a first location"; [0009], "the digital video and audio content are output in substantially in real-time"; [0038], "the stationary bike 102 may be equipped with various sensors that can measure a range of performance metrics from both the stationary bike and the rider, instantaneously"; [0039], "performance metrics that may be measured or calculated include distance, speed, resistance, power, total work, pedal cadence, heart rate"; sensing the rider's performance metrics from the stationary bike (a client device) instantaneously while the live cycling class content is output in substantially real time is obtaining, as the instructional media is being streamed, user data that pertains to performance of the activity by the user). Because Song and Foley are analogous art in the same field of endeavor of networked systems that deliver instructor-led activity instruction and customized content to users over a network, and are reasonably pertinent to the same problem of providing responsive, personalized instruction based on the user's measured performance, accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to obtain the user's performance data as the instructional media is being streamed, as taught by Foley, in the networked customized-lesson system of Song, with a reasonable expectation of success, so as to teach as the instructional media is being streamed to the client device, obtaining user data that pertains to performance of the activity by the user. This modification would have been motivated by the desire to enable the instruction to respond to the user's current, in-session performance rather than only to stored historical records and to improve the experience and provide a more engaging environment (Foley: [0006], [0073]). Song and Foley do not expressly teach wherein the computer-generated video data is generated via a video model based on image features of the instructor (interpreted as wherein, when the user-customized media includes the computer-generated video data, the computer-generated video data is generated via a video model based on image features of the instructor per the 35 U.S.C. 112(b) rejection set forth above), the video model being a first generative neural network; and the computer-generated audio data is generated via an audio model based on voice features of the instructor (interpreted as and when the user-customized media includes the computer-generated audio data, the computer-generated audio data is generated via an audio model based on voice features of the instructor per the 35 U.S.C. 112(b) rejection set forth above), the audio model being a second generative neural network. However, Savchenkov teaches wherein, when the user-customized media includes the computer-generated video data, the computer-generated video data is generated via a video model based on image features of the instructor, the video model being a first generative neural network (Savchenkov: [0065], "The sequence of the mouth texture images can be generated by a convolutional neural network (CNN)"; [0077], "a 'discriminator' neural network, such as Generative Adversarial Network (GAN), can be used with the CNN ('generator') to generate the mouth texture images"; [0068], "the CNN used in the mouth texture generation module 640 can be trained on a training set generated based on real videos recorded in a controlled environment with a single actor"; the convolutional neural network generator, trained together with a generative adversarial network on videos of a single actor (when the user-customized media includes the computer-generated video data), that generates the mouth texture images forming the target person's face (image features of the instructor) in the output video is a video model that is a first generative neural network generating computer-generated video data based on image features of the instructor); and when the user-customized media includes the computer-generated audio data, the computer-generated audio data is generated via an audio model based on voice features of the instructor, the audio model being a second generative neural network (Savchenkov: [0056], "a character embedding module 610, a deep neural network (DNN) 620"; [0058], "the DNN 620 may convert the sequence of linguistic numerical features to a sequence of sets of acoustic (numerical) features"; [0059], "Generation of acoustic numerical features can be conditioned based on speaker identification data or speaker attributes"; [0060], "the vocoder 660 ... may include a neural vocoder"; the deep neural network and neural vocoder that synthesize the audio representing the input text in the target speaker's voice (and when the user-customized media includes the computer-generated audio data), conditioned on the speaker's attributes (voice features of the instructor), are an audio model that is a second generative neural network generating computer-generated audio data based on voice features of the instructor). Because Song, in view of Foley, and Savchenkov are analogous art, with Savchenkov in the field of endeavor of computer-generated media that depicts a person and reasonably pertinent to the problem addressed by Song of providing lesson content that presents the image and voice of an instructor, accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to generate the image and voice of the instructor in the customized lesson content of Song and Foley using the convolutional-neural-network-with-generative-adversarial-network video model and the deep-neural-network-with-vocoder audio model of Savchenkov, with a reasonable expectation of success, thereby teaching wherein, when the user-customized media includes the computer-generated video data, the computer-generated video data is generated via a video model based on image features of the instructor, the video model being a first generative neural network; and when the user-customized media includes the computer-generated audio data, the computer-generated audio data is generated via an audio model based on voice features of the instructor, the audio model being a second generative neural network. This modification would have been motivated by the desire to synthesize new instructor image and voice content conveying the user-specific instruction in real time, without requiring the instructor to pre-record every possible lesson variation and to perform face reenactments that can be used in many applications (Savchenkov: [0003], [0029]). Regarding dependent claim 2, Song, in view of Foley and Savchenkov, teach the computing system of claim 1, where the instructional media includes a first portion and a second portion, where the first portion is streamed to the client device prior to the user-customized media being streamed to the client device, and further where the second portion is streamed to the client device after the user-customized media is streamed to the client device (Song: [0141], "golf practice based on the practice curriculum setting information for the user is started"; [0236], "when a preset number of user's golf shots determined to be bad has occurred, customized lesson content is output to the screen by the customized lesson provision means"; the golf lesson content presented before a bad-shot condition occurs is a first portion streamed prior to the customized lesson content, and the golf lesson content presented as the user resumes practice after the customized lesson content is output is a second portion streamed after the user-customized media). Regarding dependent claim 3, Song, in view of Foley and Savchenkov, teach the computing system of claim 1, where the user-customized media includes the computer-generated video data and the computer-generated audio data (Savchenkov: [0065], "The sequence of the mouth texture images can be generated by a convolutional neural network (CNN)"; [0095], "the method 100 may include adding, by the computing device, the audio data to the output video"; the output video containing both the generated mouth texture images of the person's face and the added synthesized audio includes both computer-generated video data and computer-generated audio data). Regarding dependent claim 4, Song, in view of Foley and Savchenkov, teach the computing system of claim 1, where the user-customized media is generated as the instructional media is being streamed to the client device (Savchenkov: [0029], "methods and systems for text and audio-based real-time face reenactment"; the instructor image and voice content are synthesized in real time, so that, in the combination in which, per Foley: [0009], Foley's class content is output in substantially real time, the user-customized media is generated as the instructional media is being streamed to the client device). Regarding dependent claim 5, Song, in view of Foley and Savchenkov, teach the computing system of claim 1, where the computer-generated video data is generated by a computer-implemented model that is trained based upon video data of the human instructor (Savchenkov: [0068], "the CNN used in the mouth texture generation module 640 can be trained on a training set generated based on real videos recorded in a controlled environment with a single actor"; [0082], "The neural network for generating the sets of facial key points can be trained on a set of real videos recorded in a controlled environment and featuring a single actor"; the convolutional neural network trained on real videos of the single actor whose face is generated is a computer-implemented model trained based upon video data of the human instructor). Regarding dependent claim 6, Song, in view of Foley and Savchenkov, teach the computing system of claim 1, where the computer-generated audio data is generated by a computer-implemented model that is trained based upon audio data that captures the voice of the human instructor (Savchenkov: [0082], "trained on a set of real videos recorded in a controlled environment and featuring a single actor speaking different predefined sentences"; [0059], "Generation of acoustic numerical features can be conditioned based on speaker identification data or speaker attributes"; the model trained on real videos of the single actor speaking, which capture that actor's voice, and conditioned on the speaker's attributes, is a computer-implemented model trained based upon audio data that captures the voice of the human instructor). Regarding dependent claim 7, Song, in view of Foley and Savchenkov, teach the computing system of claim 1, where streaming the instructional media to the client device comprises livestreaming the instructional media to the client device (Foley: [0008], "displaying information about available live and archived cycling classes that can be accessed by a first user using a first stationary bike"; [0009], "the digital video and audio content are output in substantially in real-time"; displaying a live cycling class whose video and audio content are output in substantially real time to the user's stationary bike is livestreaming the instructional media to the client device). Regarding dependent claim 8, Song, in view of Foley and Savchenkov, teach the computing system of claim 1, where the client device is a piece of exercise equipment being employed by the user to perform the activity (Foley: [0008], "a first user using a first stationary bike"; [0034], "a local system 100 comprises a stationary bike 102 with integrated or connected digital hardware including at least one display screen 104"; the stationary bike with its integrated display, on which the user performs the cycling exercise, is a client device that is a piece of exercise equipment being employed by the user to perform the activity). Regarding dependent claim 9, Song, in view of Foley and Savchenkov, teach the computing system of claim 8, where the user data comprises data output by a sensor of the exercise equipment (Foley: [0038], "the stationary bike 102 may be equipped with various sensors that can measure a range of performance metrics from both the stationary bike and the rider"; [0061], "the stationary bike 102 may be equipped with various sensors to measure and/or store data relating to user performance metrics such as speed, resistance, power, cadence, heart rate"; the performance metrics measured by the sensors of the stationary bike are user data comprising data output by a sensor of the exercise equipment). Regarding dependent claim 10, Song, in view of Foley and Savchenkov, teach the computing system of claim 1, where the user-customized media comprises the computer-generated audio data, and further where the computer-generated audio data comprises a name of the user (Savchenkov: [0035], "the computing device 110 can be configured to receive an input text 160"; [0087], "original input text 160 is replaced by the text modified by the user. A vocoder may generate an audio data to match the input text"; because the audio data is synthesized in the instructor's voice from arbitrary input text, providing the user's name as part of that input text causes the computer-generated audio data to comprise a name of the user). Regarding dependent claim 11, Song, in view of Foley and Savchenkov, teach the computing system of claim 1, where the user data comprises heart rate of the user (Foley: [0039], "performance metrics that may be measured or calculated include distance, speed, resistance, power, total work, pedal cadence, heart rate"; [0061], "various sensors to measure and/or store data relating to user performance metrics such as speed, resistance, power, cadence, heart rate"; the rider's heart rate measured by the sensors is user data comprising heart rate of the user). Regarding independent claim 12, Song teaches a method performed by a computing system, the method comprising (Song: [0011], "a user-customized practice environment provision method using virtual golf simulation, including extracting a record regarding a result of golf practice performed by a user and analyzing the record"; the recited method performed by the server and virtual golf simulation apparatuses is a method performed by a computing system): streaming instructional media simultaneously to several client devices, where the instructional media includes video of a human instructor setting forth audible instructions with respect to an activity being performed by users of the several client devices (Song: [0208], "the database 500 is preferably configured to store identification information of a plurality of virtual golf simulation apparatuses connected to the server S and user information, such as personal information of users registered in the server S"; [0239], "it is possible to provide lesson content while displaying image and voice information regarding a real lesson pro or pro golfer"; the server providing golf lesson content presenting the image and voice of a real lesson pro (video of a human instructor setting forth audible instructions) to the plurality of connected virtual golf simulation apparatuses of the registered users (several client devices) performing golf practice (an activity being performed by users of the several client devices) is streaming instructional media to several client devices); generating customized media for the user based upon the obtained user data, where the customized media for the user comprises computer-generated audiovisual data of the instructor, where the computer-generated audiovisual data pertains to the activity being performed by the user (Song: [0217], "the customized lesson provision means 300 of the virtual golf simulation apparatus generates customized lesson content to be provided to the user"; [0210], "the database 500 may store information regarding an image and a voice of a virtual lesson pro"; [0238], "it is determined that a slice golf shot has occurred in terms of the direction angle analysis item, the customized lesson provision means may extract information"; the customized lesson content, presenting both the image and the voice of the lesson pro and generated based on the analysis of the user's golf shots, is customized media for the user, comprising audiovisual data of the instructor, that pertains to the activity being performed by the user); and streaming the customized media for the user to the client device for presentment to the user as part of the instructional media being streamed to the client device while refraining from streaming the customized media to at least one other client device in the several client devices (Song: [0201], "in a so-called screen golf driving range including a plurality of virtual golf simulation apparatuses according to the present invention, a user performs golf practice through a specific virtual golf simulation apparatus"; [0215], "transmit the extracted information to the virtual golf simulation apparatus"; [0217], "generates customized lesson content to be provided to the user"; because the customized lesson content is generated for, and transmitted to, the specific apparatus of the user whose golf shots were analyzed, and differs from the content generated for other registered users, streaming that content to that user's apparatus is done while refraining from streaming it to at least one other client device among the several client devices). Song does not expressly teach as the instructional media is being streamed to the several client devices, obtaining user data from a client device from amongst the several client devices, where the user data pertains to performance of the activity by a user of the client device. However, Foley teaches as the instructional media is being streamed to the several client devices, obtaining user data from a client device from amongst the several client devices, where the user data pertains to performance of the activity by a user of the client device (Foley: [0009], "the digital video and audio content are output in substantially in real-time"; [0038], "the stationary bike 102 may be equipped with various sensors that can measure a range of performance metrics from both the stationary bike and the rider, instantaneously"; [0077]-[0078], "the networked exercise system may be configured with a plurality of user bikes 400 in communication"; sensing, from a particular stationary bike among the plurality of user bikes, that rider's performance metrics while the class content is output in substantially real time is obtaining, as the instructional media is being streamed to the several client devices, user data from a client device from amongst the several client devices that pertains to that user's performance). Song and Foley are analogous art in the same field of endeavor of networked systems that deliver instructor-led activity instruction to multiple users over a network, and are reasonably pertinent to the same problem of tailoring the instruction to an individual user's measured performance, accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to obtain, as the instructional media is being streamed to the several client devices, the performance data of a user from that user's client device, as taught by Foley, in the networked customized-lesson method of Song, with a reasonable expectation of success, so as to teach as the instructional media is being streamed to the several client devices, obtaining user data from a client device from amongst the several client devices, where the user data pertains to performance of the activity by a user of the client device. This modification would have been motivated by the desire to enable the instruction to respond to the user's current, in-session performance rather than only to stored historical records and to improve the experience and provide a more engaging environment (Foley: [0006], [0073]). Song and Foley do not expressly teach wherein the computer-generated audiovisual data is generated via a multimodal model based on audiovisual features of the instructor (interpreted as generated via an audio model and a video model based on an audiovisual data of the instructor per the 35 U.S.C. 112(a) rejection above). However, Savchenkov teaches wherein the computer-generated audiovisual data is generated via an audio model and a video model based on an audiovisual data of the instructor (Savchenkov: [0065], "The sequence of the mouth texture images can be generated by a convolutional neural network (CNN)"; [0077], "a 'discriminator' neural network, such as Generative Adversarial Network (GAN), can be used with the CNN ('generator')"; [0058], "the DNN 620 may convert the sequence of linguistic numerical features to a sequence of sets of acoustic (numerical) features"; [0060], "the vocoder 660 ... may include a neural vocoder"; [0095], "the method 100 may include adding, by the computing device, the audio data to the output video"; the combination of the convolutional-neural-network-with-generative-adversarial-network video model and the deep-neural-network-with-vocoder audio model, which produces an output video of the person's face with the synthesized audio of the person's voice added). Because Song, in view of Foley, and Savchenkov are analogous art, with Savchenkov in the field of endeavor of computer-generated media that depicts a person and reasonably pertinent to the problem addressed by Song of providing lesson content that presents the image and voice of an instructor, accordingly, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to generate the audiovisual customized lesson content of Song, as modified by Foley, using the convolutional-neural-network-with-generative-adversarial-network video model together with the deep-neural-network-with-vocoder audio model of Savchenkov, with a reasonable expectation of success, thereby teaching wherein the computer-generated audiovisual data is generated via an audio model and a video model based on an audiovisual data of the instructor. This modification would have been motivated by the desire to synthesize new instructor image and voice content conveying the user-specific instruction in real time, without requiring the instructor to pre-record every possible lesson variation and to perform face reenactments that can be used in many applications (Savchenkov: [0003], [0029]). Regarding dependent claim 13, Song, in view of Foley and Savchenkov, teach the method of claim 12, where the audiovisual data of the instruct[[ion]]or comprises computer-generated images of the instructor and computer-generated audio in a voice of the instructor (Savchenkov: [0065], "The sequence of the mouth texture images can be generated by a convolutional neural network (CNN)"; [0060], "the vocoder 660 ... may include a neural vocoder"; [0095], "adding, by the computing device, the audio data to the output video"; the generated mouth texture images forming the person's face and the synthesized audio in the person's voice added to the output video are computer-generated images of the instructor and computer-generated audio in a voice of the instructor). Regarding dependent claim 14, Song, in view of Foley and Savchenkov, teach the method of claim 13, where the computer-generated images of the instructor are generated by a computer-implemented model that has been trained based upon video of the instructor (Savchenkov: [0068], "the CNN used in the mouth texture generation module 640 can be trained on a training set generated based on real videos recorded in a controlled environment with a single actor"; [0082], "trained on a set of real videos recorded in a controlled environment and featuring a single actor"; the convolutional neural network trained on real videos of the single actor whose face is generated is a computer-implemented model trained based upon video of the instructor). Regarding dependent claim 15, Song, in view of Foley and Savchenkov, teach the method of claim 12, further comprising: as the instructional media is being streamed to the several client devices including a first client device and a second client device, obtaining second user data from the second client device from amongst the several client devices, where the second user data pertains to performance of the activity by a second user of the second client device (Foley: [0038], "various sensors that can measure a range of performance metrics from both the stationary bike and the rider, instantaneously"; [0077]-[0078], "a plurality of user bikes 400 in communication"; sensing a second rider's performance from that rider's stationary bike among the plurality of bikes obtains second user data from the second client device), generating second customized media for the second user based upon the obtained second user data (Song: [0217], "generates customized lesson content to be provided to the user"; generating customized lesson content for the second user based on that user's analyzed shots is generating second customized media for the second user), and streaming the second customized media for the second user to the second client device for presentment to the second user as part of the instructional media being streamed to the second client device while refraining from streaming the second customized media to the at least one other client device in the several client devices (Song: [0201], "a user performs golf practice through a specific virtual golf simulation apparatus"; [0215], "transmit the extracted information to the virtual golf simulation apparatus"; because the second user's customized lesson content is generated for, and transmitted to, that user's specific apparatus and differs from the content for other users, it is streamed to the second client device while refraining from streaming it to the at least one other client device among the several client devices). Regarding dependent claim 16, Song, in view of Foley and Savchenkov, teach the method of claim 15, where the customized media for the user and the second customized media for the second user are streamed to the first client device and the second client device, respectively, simultaneously (Foley: [0077], "the system can provide for simultaneous participation by multiple users in a recorded class, synchronized by the system"; [0078], "the networked exercise system may be configured with a plurality of user bikes 400"; streaming the respective customized content to the first and second users' bikes as they participate simultaneously in the synchronized class streams the customized media and the second customized media to the first and second client devices, respectively, simultaneously). Regarding dependent claim 17-19, these claims are substantially the same as the computing system claims 2&7-8, respectively. Thus, claims 17-19 are rejected for the same reasons as claims 2&7-8. Regarding independent claim 20, it is a computer-readable storage medium claim that is substantially the same as the computing system of claim 1. Thus claim 20 is rejected for the same reason as claim 1. In addition Song teaches a computer-readable storage medium comprising instructions that, when executed by a processor, cause the processor to perform acts comprising (Song: [0082]). Response to Arguments Applicant’s SPEC amendments and Remarks filed 5/5/2026 with respect to the SPEC objections set forth in the Office Action dated 12/5/2025 are persuasive and thus the said SPEC objections are withdrawn. Applicant’s filing of a terminal disclaimer referencing US Patent No. 11,771,977 was accepted and thus the nonstatutory double patenting rejection set forth in the Office Action dated 12/5/2025 is withdrawn. Applicant’s claim amendments and Remarks filed 5/5/2026 with respect to the 35 U.S.C. 103 rejections set forth in the Office Action dated 12/5/2025 are not persuasive. In response to Applicant's amendment, claims 1 through 20 stand rejected under 35 U.S.C. 103 over Song in view of Foley and further in view of Savchenkov, as set forth above; these grounds were necessitated by Applicant's amendment. Applicant's arguments directed to the amended first and second generative neural network limitations and to Song are addressed on the merits below. Applicant's arguments directed to Russell, Theis, and Yoo, to the previously applied four-reference combination, and to the nonstatutory double patenting rejection are moot in view of the new grounds of rejection, as detailed above. Applicant argues that Song does not disclose the generative neural network configuration and that such generative neural network models had not been developed when Song was filed in 2012. The Examiner respectfully disagrees. The date on which Song was filed is immaterial to the rejection. Obviousness is assessed as of the effective filing date of the claimed invention (May 20, 2021), not the filing date of any single applied reference. The generative-neural-network limitation is supplied by Savchenkov, published July 23, 2020, which is prior art under 35 U.S.C. 102(a)(1). That generative neural networks postdate Song's 2012 filing does not remove Savchenkov from the prior art or negate the obviousness of the Song, Foley, and Savchenkov combination. Accordingly, the argument is not persuasive and the rejection is maintained. Applicant requests clarification of the status of claims 15 and 16, which were not addressed in the Office Action dated December 5, 2025. In this Office Action, claims 1 through 20, including claims 15 and 16, are rejected under 35 U.S.C. 103 over Song in view of Foley and further in view of Savchenkov. Applicant's request for clarification is thereby addressed. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KUANG FU CHEN whose telephone number is (571)272-1393. The examiner can normally be reached M-F 9:00-5:30pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached on (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KC CHEN/Primary Patent Examiner, Art Unit 2143
Read full office action

Prosecution Timeline

Sep 01, 2023
Application Filed
Dec 05, 2025
Non-Final Rejection mailed — §103, §112
May 05, 2026
Response Filed
Jul 07, 2026
Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12682047
Machine Learning Time Series Anomaly Detection
4y 1m to grant Granted Jul 14, 2026
Patent 12675771
System for Online Interaction with Content
5y 8m to grant Granted Jul 07, 2026
Patent 12664448
AVERAGE TREATMENT EFFECT FOR PAIRED DATA
4y 11m to grant Granted Jun 23, 2026
Patent 12657260
SIMULATING TRAINING DATA TO MITIGATE BIASES IN MACHINE LEARNING MODELS
4y 1m to grant Granted Jun 16, 2026
Patent 12657494
LEARNING SYSTEM, LEARNING METHOD, AND STORAGE MEDIUM
2y 12m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
80%
Grant Probability
99%
With Interview (+68.4%)
2y 11m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 270 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month