Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 05/22/2026 was filed before the mailing date of this office action. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Status of Claims
This action is in response to the amendments filed 05/22/2026. Claims 1, 5-6, 8-10, and 14-17 have been amended, claims 7, 13, and 18-20 have been cancelled, claims 21-25 have been added. Claims 1-6, 8-12, 14-17, and 21-25 are currently pending.
Response to Arguments
Claim 7, 13, and 18-20 have been cancelled, therefore the rejection of claims 7, 13, and 18-20 no longer stand.
In light of Applicant’s amendment, the 112(b) rejection of claim 15 has been withdrawn.
Applicant’s arguments regarding the prior art rejection have been fully considered but they are not persuasive. With regards to claims 1 and 25, Applicant argues that the Hofshi reference “merely modifies existing media material in real-time based on a determined emotional state without generating a script, let along constructing a modality script comprising a plurality of time-segmented directives” and that Hofshi lacks any suggestion to “predefine a structured set of time-segmented directives” in an “orchestration layer”. Examiner respectfully disagrees and notes that the claim does not recite “predefining” a structured set of time-segmented directives or an “orchestration layer”. Examiner notes that Applicant’s original disclosure does not recite definitions for a “modality script” or “directives” such that the broadest reasonable interpretation of a modality script would not include the stream, or script, of time-segmented audiovisual frames which are arranged to elicit a desired emotional response, or directive, from a user based on detected emotion data.
With regards to claim 10, Applicant argues that the prior art does not teach the limitation directed to assigning weights to dimensions of the emotion state vector “according to a weighting profile defined by a user-selected template”. Examiner notes that this limitation has been rejected under 35 U.S.C. 112(a) as lacking support from Applicant’s original disclosure. Given the interpretation of this limitation from the 112(a) rejection, Examiner notes that at least paragraphs [0027]-[0028] of the Hofshi reference teach that different dimensions, or regions, of the ESV are given different values, or weights, depending on the user’s state, or use case.
Lastly, Examiner notes that the Amores Fernandez reference has been brought in to teach the virtual reality headset in claim 17. The prior art rejections have been updated to include the amended limitations and to clarify the reasoning given for the limitations that were not amended where necessary.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 5, 10-12, 14-17, and 24 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claim 5 recites the limitation “wherein the dimensions of the ESV (i.e., emotion state vector) are each given a respective weight according to a weighting profile defined by a user-selected template”. Examiner notes that on page 1 of Applicant’s remarks, Applicant has pointed to page 21 lines 25-27 and page 22 lines 1-2 to provide support for the amendment to claim 5; however, these sections merely state that “Dimensions can be weighted by a weight vector. . .configured to tune relative importance of the dimensions depending on the use case”. Examiner further notes that Applicant’s original disclosure also teaches that a modality script generator can support user-editable templates related to “dynamic switching between segment variants” on at least page 6 and “breakout/breakdown interludes, guided-meditation flows, mnemonic explainers, and content magnification transitions” on at least page 15, but Applicant’s original disclosure does not support wherein these user-editable templates are used for weighting dimensions of the ESV. For purposes of examination, Examiner is interpreting that dimensions in the ESV can be given different weights depending on a selected use case, as supported by page 21 lines 25-27 and page 22 lines 1-2 of Applicant’s specification and claim 5 as originally filed.
Claim 10 recites a similar limitation and is rejected for the same reasons. Dependent claims 11-12, 14-17, and 24 are also rejected because they fail to correct the deficiencies of claim 10 on which they depend.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 4-6, 8-11, 14-16, and 21-25 are rejected under 35 U.S.C. 103 as being unpatentable over Hofshi (US 20120194648 A1, herein Hofshi) in view of Labbé et al (US 20230113072 A1, herein Labbé).
Regarding claim 1, Hofshi teaches a computer-implemented method (para. [0002] recites “Embodiments of the invention relate to methods and devices for modifying video and/or audio material in real time”), comprising: automatically constructing an emotion state vector (ESV) configured to encode an affective state of a user, wherein the ESV comprises a plurality of dimensions each based on user-specific data (para. [0005]-[0007] recite “The apparatus includes a controller having a processor for processing the physiological data to determine an emotional state of the user, and for modifying the V/A material in real time responsive to the determined emotional state. In an embodiment of the invention, the controller processes the received physiological data to generate a vector, hereinafter referred to as an emotion state vector (ESV), which comprises a plurality of components and provides a measure of an emotional state of the user. In an embodiment, the plurality of components comprises an arousal component and an attitude component”. Para. [0024] recites “Determining the emotional state of the user may include determining one or more emotional parameters of the user. For example, arousal, anxiety, relaxation, apathy, etc. According to an embodiment, the emotional state is determined based on at least two emotional parameters, for example an arousal level and an attitude. The arousal level represents an intensity or challenge that the user feels while engaging with V/A material, and attitude represents a general attitude and sentiment of the user toward the V/A material with which he or she is engaging.” (i.e., constructing an emotion state vector, or ESV, comprising different dimensions related to the emotional state of a user, or user-specific data, based on sensed input data));
constructing, based on the ESV, a modality script comprising a plurality of time-segmented directives including information defining one or more respective media content items; arranging the media content items into an audio or audiovisual media output according to the modality script (para. [0010] recites “For each V/A data frame, the emotion interface apparatus determines an emotion state vector (ESV) of the user and locates a region, hereinafter an emotion focal region (EFR), in emotion space to which the ESV points”. Para. [0030] recites “V/A material is associated with an emotion trajectory that defines an emotion profile for the V/A material and EI (i.e., emotion interface) apparatus 102 modifies presentation, or streaming, of the V/A material in accordance with the user's EFR (i.e., emotional focal region). Optionally, the emotion trajectory defines expected or normative emotional states for a user of the V/A material. Optionally, the emotion trajectory defines emotional states that are extreme and/or highly undesirable”. Para. [0032] recites “EI apparatus 102 modifies streaming of V/A material responsive to a relationship between the user's EFR generated at a time when the user is presented with the V/A frame and the FEZ (i.e., frame emotion zone) for the V/A frame provided by the emotion trajectory”. Para. [0034] recites “EI apparatus 102 controls streaming of V/A frames 120 responsive to a relationship between FEZ 117 and EFR 115. Optionally, controller 107 determines a distance between EFR 115 and FEZ 117 and controls streaming of V/A material 119 responsive to the determined distance” (i.e., constructing and arranging a series of time-segmented audiovisual media content frames, or modality script, based on information from the ESV));
and providing the media output to enable playback on an electronic device of the user (para. [0029] recites “In response to the user's emotional state, EI apparatus 102 modifies the progression, or streaming, of V/A material provided to user 106 by computer 105 in real time. In a situation for which the V/A material is a video game, if EI apparatus 102 receives indication that user 106 is distressed, the apparatus may modify the level of difficulty of the game, so as to provide the user with a material which will lower indications of user distress. In a situation for which the V/A material is a movie, EI apparatus 102 may modify the movie in response to the emotional state of user 106, by displaying a progression of scenes that increases the user's sense of relaxation” (i.e., providing the arranged media output to the user device)).
However, Hofshi does not explicitly teach generating at least a portion of the one or more media content items for each time-segmented directive using one or more generative artificial intelligence (AI) engines.
Labbé teaches generating at least a portion of the one or more media content items for each time-segmented directive using one or more generative artificial intelligence (AI) engines (para. [0216] recites “The MIR generator process 1900 is used to generate a MIR blueprint for an audio segment (e.g., a song) that is intended to induce a specific affective response in listeners. The MIR blueprint generated by the MIR generator process 1900 typically identifies MIR features of the song as a whole, as well as MIR features of each of multiple epochs (i.e., time-segmented features) of the audio segment, that will induce the desired affective response”. Para. [0225] recites “The functionality of the MIR generation process 1900 can also be achieved by other generative deep learning modalities like variational autoencoders (VAE) or simply a recurrent neural network (RNN) on its own. The GAN model has been evaluated as an effective means to execute the needed functionality but additional, similar algorithms may also be effective, particularly as advances in machine learning occur” (i.e., using a generating artificial intelligence engine, or model, to generate media content)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine these teachings by utilizing the generative artificial intelligence model from Labbé to generate personalized content for the system from Hofshi to better adapt to a user’s determined emotional state. Hofshi and Labbé are both directed to systems which personalize media content based on a user’s determined emotional state. One of ordinary skill would recognize that a generative artificial intelligence model, like the one taught by Labbé, could generate media content that is more closely tailored to the user’s current emotional state determined by the system from Hofshi.
Regarding claim 2, the combination of Hofshi and Labbé teaches the method of claim 1 as mentioned above, further comprising: receiving a baseline personalization profile (BPP) based on stated preferences of the user; wherein the modality script is further based on the BPP (Hofshi para. [0023]-[0024] recite “the physiological measurements of a user may be collected over time for determining a user profile. The user profile may include an expected range of values of a physiological parameter for that user. In addition, the user profile may include patterns of values of a physiological parameter of the user, characterizing the user's emotional response to the V/A material. Determining the emotional state of the user may include determining one or more emotional parameters of the user. For example, arousal, anxiety, relaxation, apathy, etc.”. Hofshi para. [0025] recites “when only the attitude is positive but the arousal level is low, the user may like the V/A material, but he or she might not be challenged enough, and may be bored. Similarly, when the arousal level is high and the attitude is negative, whereas the user is challenged, the user might be tense, since the material may be too difficult for him or her. In an embodiment, the V/A material may be adjusted by EI apparatus 102 to simultaneously provide a satisfactory arousal level and positive attitude for the user” (i.e., collecting user preferences to create a personalized profile, which can be used to determine when to script, or change the media stream)).
Regarding claim 4, the combination of Hofshi and Labbé teaches the method of claim 1 as mentioned above, wherein the dimensions of the ESV include one or more dimensions selected from the list consisting of arousal, valence, focus, fatigue, confidence, and readiness (Hofshi para. [0023]-[0024] recite “the user profile may include patterns of values of a physiological parameter of the user, characterizing the user's emotional response to the V/A material. Determining the emotional state of the user may include determining one or more emotional parameters of the user. For example, arousal, anxiety, relaxation, apathy, etc. The emotional state is determined based on at least two emotional parameters, for example an arousal level and an attitude. The arousal level represents an intensity or challenge that the user feels while engaging with V/A material, and attitude represents a general attitude and sentiment of the user toward the V/A material with which he or she is engaging”. Hofshi para. [0027] recites “The FIGURE shows an ESV 108 defined in the emotion space for user 106. By way of example, emotion space 110 includes a plurality of regions 116 schematically shown in space 110 that are identified with emotional states, such as, anxiety, arousal, worry, flow, apathy, boredom, etc.” (i.e., the dimensions of the ESV can include at least arousal and determinations of positive or negative attitude, or valence)).
Regarding claim 5, the combination of Hofshi and Labbé teaches the method of claim 1 as mentioned above, wherein the dimensions of the ESV are each given a respective weight according to a weighting profile defined by a user-selected template (Hofshi para. [0027]-[0028] recite “FIG. 1, schematically shows an emotion space 110 having an arousal axis 112 and an attitude axis 114. The FIGURE shows an ESV 108 defined in the emotion space for user 106. By way of example, emotion space 110 includes a plurality of regions 116 schematically shown in space 110 that are identified with emotional states, such as, anxiety, arousal, worry, flow, apathy, boredom, etc. ESV 108 points to an emotional focal region (EFR) 115 in emotion space 110, representing a region in emotion space that is descriptive of an emotion state of user 106. A size of region EFR 115 is indicative of a variance for the emotion state. EFR 115 may be located inside one of emotional regions 116, or straddle two adjacent regions 116. In FIG. 1, for the given configuration of emotion space, EFR 115 indicates a state of anxiety for user 106. Were the user simultaneously exhibiting strong arousal and strong satisfaction, EFR 115 would have been located in the upper right hand region of emotion space 110 and the user would be considered to be in an emotional state of "flow", conventionally also referred to as "being in the zone"” (Examiner notes that this limitation is interpreted in light of the 112(a) rejection such that the dimensions of the ESV are given different values, or weights, depending on the user’s current state, or use case. Given this interpretation, Hofshi teaches that different dimensions, or regions, of the ESV are given different values, or weights, depending on the user’s state, or use case)).
Regarding claim 6, the combination of Hofshi and Labbé teaches the method of claim 1 as mentioned above, further comprising: normalizing the media content items before arranging the media content items (Labbé para. [0182] recites “a user who frequently feels sad but rarely feels energetic may have his or her affective inference neural network 140 calibrated to normalize the weights given to these states based on a baseline or average set of affective state values specific to the user. The system may also use this user input data to make recommendations to the user for how to employ the system to achieve the user's goals, such as mental health or mood management goals.” (i.e., media content can be normalized before it is arranged for the user”)).
Regarding claim 8, the combination of Hofshi and Labbé teaches the method of claim 1 as mentioned above, further comprising: dynamically refreshing the ESV in response to receiving new user-specific data (Hofshi para. [0029] recites “In response to the user's emotional state, EI apparatus 102 modifies the progression, or streaming, of V/A material provided to user 106 by computer 105 in real time. For example, in a situation for which the V/A material is a video game, if EI apparatus 102 receives indication that user 106 is distressed, the apparatus may modify the level of difficulty of the game, so as to provide the user with a material which will lower indications of user distress. By way of another example, in a situation for which the V/A material is a movie, EI apparatus 102 may modify the movie in response to the emotional state of user 106, by displaying a progression of scenes that increases the user's sense of relaxation” (i.e., refreshing the ESV as the user-specific sensed inputs are received in real time, or dynamically)).
Regarding claim 9, the combination of Hofshi and Labbé teaches the method of claim 1 as mentioned above, wherein at least one dimension of the ESV is further determined based on a sensed biometric input from a wearable device (Hofshi para. [0019] recites “Bracelet sensor 104, comprises at least one device suitable for measuring a physiological parameter of user 106 usable to determine an emotional state of the user. Optionally, bracelet sensor 104 comprises at least one of an ECG device for measuring heart rate, a thermometer for measuring skin temperature, an acoustic detector for measuring vascular activity, or any other sensor for measuring physiological parameters”. Hoshi para. [0022] recites “Controller 107 processes the physiological measurements it receives from bracelet sensor 104 and images it receives from imaging system 150 to determine an emotional state of user 106 using any known method of inferring an emotional state responsive to changes in status of human physiology” (i.e., determining aspects of the ESV based on biometric inputs from a wearable device)).
Regarding claim 10, Hofshi teaches a system for generating personalized digital content (para. [0004] recites “An embodiment of the invention provides an emotion interface apparatus, that operates to interface emotions of a person, hereinafter also a "user", interacting with video and/or audio (video/audio) material, and modifies the video/audio (V/A) material in real time responsive to the emotions”), the system comprising: one or more processors; a memory; and a plurality of instructions stored in the memory, wherein the plurality of instructions, when executed by the one or more processors, are configured to:
automatically construct an emotion state vector (ESV) configured to encode an affective state of a user, wherein the ESV comprises a plurality of dimensions each based on user-specific data a respective sensed input related to the user (para. [0005]-[0007] recite “The apparatus includes a controller having a processor for processing the physiological data to determine an emotional state of the user, and for modifying the V/A material in real time responsive to the determined emotional state. In an embodiment of the invention, the controller processes the received physiological data to generate a vector, hereinafter referred to as an emotion state vector (ESV), which comprises a plurality of components and provides a measure of an emotional state of the user. In an embodiment, the plurality of components comprises an arousal component and an attitude component”. Para. [0024] recites “Determining the emotional state of the user may include determining one or more emotional parameters of the user. For example, arousal, anxiety, relaxation, apathy, etc. According to an embodiment, the emotional state is determined based on at least two emotional parameters, for example an arousal level and an attitude. The arousal level represents an intensity or challenge that the user feels while engaging with V/A material, and attitude represents a general attitude and sentiment of the user toward the V/A material with which he or she is engaging.” (i.e., constructing an emotion state vector, or ESV, comprising different dimensions related to the emotional state of a user, or user-specific data, based on sensed input data));
construct, based on the ESV, a modality script comprising information defining and orchestrating a plurality of media content items (para. [0010] recites “For each V/A data frame, the emotion interface apparatus determines an emotion state vector (ESV) of the user and locates a region, hereinafter an emotion focal region (EFR), in emotion space to which the ESV points”. Para. [0030] recites “V/A material is associated with an emotion trajectory that defines an emotion profile for the V/A material and EI (i.e., emotion interface) apparatus 102 modifies presentation, or streaming, of the V/A material in accordance with the user's EFR (i.e., emotional focal region). Optionally, the emotion trajectory defines expected or normative emotional states for a user of the V/A material. Optionally, the emotion trajectory defines emotional states that are extreme and/or highly undesirable” (i.e., constructing series of time-segmented audiovisual media content frames, or modality script, based on information from the ESV));
assign a weight to each of the plurality of dimensions of the ESV according to a weighting profile defined by a user-selected template (Hofshi para. [0027]-[0028] recite “FIG. 1, schematically shows an emotion space 110 having an arousal axis 112 and an attitude axis 114. The FIGURE shows an ESV 108 defined in the emotion space for user 106. By way of example, emotion space 110 includes a plurality of regions 116 schematically shown in space 110 that are identified with emotional states, such as, anxiety, arousal, worry, flow, apathy, boredom, etc. ESV 108 points to an emotional focal region (EFR) 115 in emotion space 110, representing a region in emotion space that is descriptive of an emotion state of user 106. A size of region EFR 115 is indicative of a variance for the emotion state. EFR 115 may be located inside one of emotional regions 116, or straddle two adjacent regions 116. In FIG. 1, for the given configuration of emotion space, EFR 115 indicates a state of anxiety for user 106. Were the user simultaneously exhibiting strong arousal and strong satisfaction, EFR 115 would have been located in the upper right hand region of emotion space 110 and the user would be considered to be in an emotional state of "flow", conventionally also referred to as "being in the zone"” (Examiner notes that this limitation is interpreted in light of the 112(a) rejection such that the dimensions of the ESV are given different values, or weights, depending on the user’s current state, or use case. Given this interpretation, Hofshi teaches that different dimensions, or regions, of the ESV are given different values, or weights, depending on the user’s state, or use case));
arrange the plurality of media content items into an audio or audiovisual media output stream according to the modality script (Para. [0032] recites “EI apparatus 102 modifies streaming of V/A material responsive to a relationship between the user's EFR generated at a time when the user is presented with the V/A frame and the FEZ for the V/A frame provided by the emotion trajectory”. Para. [0034] recites “EI apparatus 102 controls streaming of V/A frames 120 responsive to a relationship between FEZ 117 and EFR 115. Optionally, controller 107 determines a distance between EFR 115 and FEZ 117 and controls streaming of V/A material 119 responsive to the determined distance” (i.e., arranging the constructed series of time-segmented audiovisual media content frames, or modality script, based on information from the ESV)); and
provide the media output to enable playback on an electronic device of the user (para. [0029] recites “In response to the user's emotional state, EI apparatus 102 modifies the progression, or streaming, of V/A material provided to user 106 by computer 105 in real time. In a situation for which the V/A material is a video game, if EI apparatus 102 receives indication that user 106 is distressed, the apparatus may modify the level of difficulty of the game, so as to provide the user with a material which will lower indications of user distress. In a situation for which the V/A material is a movie, EI apparatus 102 may modify the movie in response to the emotional state of user 106, by displaying a progression of scenes that increases the user's sense of relaxation” (i.e., providing the arranged media output to the user device)).
However, Hofshi does not explicitly teach generating at least a portion of the plurality of media content items by communicating with one or more generative artificial intelligence (Al) engines.
Labbé teaches generating at least a portion of the plurality of media content items by communicating with one or more generative artificial intelligence (Al) engines (para. [0216] recites “The MIR generator process 1900 is used to generate a MIR blueprint for an audio segment (e.g., a song) that is intended to induce a specific affective response in listeners. The MIR blueprint generated by the MIR generator process 1900 typically identifies MIR features of the song as a whole, as well as MIR features of each of multiple epochs (i.e., time-segmented features) of the audio segment, that will induce the desired affective response”. Para. [0225] recites “The functionality of the MIR generation process 1900 can also be achieved by other generative deep learning modalities like variational autoencoders (VAE) or simply a recurrent neural network (RNN) on its own. The GAN model has been evaluated as an effective means to execute the needed functionality but additional, similar algorithms may also be effective, particularly as advances in machine learning occur” (i.e., using a generating artificial intelligence engine, or model, to generate media content)).
See claim 1 for motivation to combine.
Claim 11 is a system claim and its limitation is included in claim 2. Claim 11 is rejected for the same reasons as claim 2.
Claim 14 is a system claim and its limitation is included in claim 6. Claim 14 is rejected for the same reasons as claim 6.
Regarding claim 15, the combination of Hofshi and Labbé teaches the system of claim 14 as mentioned above, wherein arranging the plurality of media content items into the media output comprises synchronizing the normalized plurality of media content items to produce a playable media output comprising aligned media content items (Labbé para. [0256] recites “The required MIR data 1757 is broken down into epochs (i.e. time periods) of MIR data corresponding to the MIR features needed for each epoch of the mastered audio segment by a MIR epoch splitting process 2114. These epochs of MIR data are referred to as target MIR epochs 2116, indicating the MIR feature targets for the mastering process for a given epoch. The epoch sizes are synchronized between the epoch splitting process 2106 and MIR epoch splitting process 2114 in order to maintain the same timeline throughout the mastering process” (i.e., the normalized generated media content from at least paragraph [0182] of Labbé can be synchronized to produce playable media)).
Claim 16 is a system claim and its limitation is included in claim 8. Claim 16 is rejected for the same reasons as claim 8.
Regarding claim 21, the combination of Hofshi and Labbé teaches the method of claim 10 as mentioned above, wherein the information for each of the plurality of time-segmented directives comprises a segment role, a framing method, a media modality type, a tone, a voice, a prosody, a genre, a time duration, a tempo, a persona, an emotional framing, or a pacing (Hofshi para. [0032] recites “EI apparatus 102 modifies streaming of V/A material responsive to a relationship between the user's EFR (i.e., emotional focal region) generated at a time when the user is presented with the V/A frame and the FEZ (i.e., frame emotion zone) for the V/A frame provided by the emotion trajectory”. Hofshi para. [0034] recites “EI apparatus 102 controls streaming of V/A frames 120 responsive to a relationship between FEZ 117 and EFR 115. Optionally, controller 107 determines a distance between EFR 115 and FEZ 117 and controls streaming of V/A material 119 responsive to the determined distance” (i.e., determining an emotional framing based on the emotional focal region and frame emotion zone to determine if the media content for a given time-segment should be changed)).
Regarding claim 22, the combination of Hofshi and Labbé teaches the method of claim 1 as mentioned above, wherein at least two of the plurality of time-segmented directives include synchronization information defining at least partial temporal alignment between the corresponding media content items (Labbé para. [0217] recites “The MIR generator process 1900 is used to generate a MIR blueprint for an audio segment (e.g., a song) that is intended to induce a specific affective response in listeners. The MIR blueprint generated by the MIR generator process 1900 typically identifies MIR features of the song as a whole, as well as MIR features of each of multiple epochs (i.e., temporal sub-segments) of the audio segment, that will induce the desired affective response”. Labbé para. [0256] recites “The required MIR data 1757 is broken down into epochs (i.e. time periods) of MIR data corresponding to the MIR features needed for each epoch of the mastered audio segment by a MIR epoch splitting process 2114. These epochs of MIR data are referred to as target MIR epochs 2116, indicating the MIR feature targets for the mastering process for a given epoch. The epoch sizes are synchronized between the epoch splitting process 2106 and MIR epoch splitting process 2114 in order to maintain the same timeline throughout the mastering process” (i.e., media content items from at least paragraph [0182] of Labbé can be synchronized, or temporally aligned, over multiple time-segments, or epochs, to produce playable media)).
Regarding claim 23, the combination of Hofshi and Labbé teaches the method of claim 1 as mentioned above, wherein the user-specific data comprises one or more of user goals, user challenges, user objectives, conversations with an Al coach, structured prompts, surveys, assessments, journal entries, user-provided material, user preferences, profile data, or contextual metadata (Hofshi para. [0023] recites “the physiological measurements of a user may be collected over time for determining a user profile. The user profile may include an expected range of values of a physiological parameter for that user. In addition, the user profile may include patterns of values of a physiological parameter of the user, characterizing the user's emotional response to the V/A material” (i.e., user-specific data can include at least profile data)).
Regarding claim 24, the combination of Hofshi and Labbé teaches the system of claim 10 as mentioned above, wherein the modality script comprises a plurality of time-coded segments, wherein each segment corresponds to one or more media content items of the plurality of media content items and includes directives configured to control at least partial generation of the one or more media content items by the one or more generative Al engines (Hofshi para. [0010] recites “For each V/A data frame, the emotion interface apparatus determines an emotion state vector (ESV) of the user and locates a region, hereinafter an emotion focal region (EFR), in emotion space to which the ESV points”. Hofshi para. [0030] recites “the emotion trajectory defines expected or normative emotional states for a user of the V/A material. Optionally, the emotion trajectory defines emotional states that are extreme and/or highly undesirable”. Hofshi para. [0032] recites “EI apparatus 102 modifies streaming of V/A material responsive to a relationship between the user's EFR generated at a time when the user is presented with the V/A frame and the FEZ for the V/A frame provided by the emotion trajectory”. Hofshi para. [0034] recites “EI apparatus 102 controls streaming of V/A frames 120 responsive to a relationship between FEZ 117 and EFR 115. Optionally, controller 107 determines a distance between EFR 115 and FEZ 117 and controls streaming of V/A material 119 responsive to the determined distance”. Labbé para. [0217] recites “The MIR generator process 1900 is used to generate a MIR blueprint for an audio segment (e.g., a song) that is intended to induce a specific affective response in listeners. The MIR blueprint generated by the MIR generator process 1900 typically identifies MIR features of the song as a whole, as well as MIR features of each of multiple epochs (i.e., temporal sub-segments) of the audio segment, that will induce the desired affective response”. Labbé para. [0225] recites “The functionality of the MIR generation process 1900 can also be achieved by other generative deep learning modalities like variational autoencoders (VAE) or simply a recurrent neural network (RNN) on its own. The GAN model has been evaluated as an effective means to execute the needed functionality but additional, similar algorithms may also be effective, particularly as advances in machine learning occur” (i.e., a modality stream, or script, of time-coded segments, or frames, of audiovisual media content can be orchestrated and can be modified by a generative model to produce a desired emotional state in a user)).
Regarding claim 25, Hofshi teaches a computer-implemented method (para. [0002] recites “Embodiments of the invention relate to methods and devices for modifying video and/or audio material in real time”), comprising: determining a user objective based on user-specific data (para. [0015]-[0016] recites “If the EI apparatus receives an indication that the user is bored with the music, the EI apparatus may control the audio device to modify the music so as to improve the user's interest in the music being played. By way of another example, an EI apparatus in accordance with an embodiment of the invention may be coupled to a game console interfacing a user with a video game, and may modify the level of difficulty of challenges with which the video game challenges the user responsive to the user's emotional state. If the EI apparatus receives an indication that the user is negatively or unduly stressed by the game, the apparatus may change the level of difficulty of the game” (i.e., determining a user objective, such as alleviating boredom or passing a difficult video game challenge, based on user-specific emotion data));
constructing, based on the determined user objective, a modality script comprising a plurality of time-segmented directives each including information defining one or more respective media content items; arranging the media content items into an audio or audiovisual media output according to the modality script (para. [0010] recites “For each V/A data frame, the emotion interface apparatus determines an emotion state vector (ESV) of the user and locates a region, hereinafter an emotion focal region (EFR), in emotion space to which the ESV points”. Para. [0030] recites “V/A material is associated with an emotion trajectory that defines an emotion profile for the V/A material and EI (i.e., emotion interface) apparatus 102 modifies presentation, or streaming, of the V/A material in accordance with the user's EFR (i.e., emotional focal region). Optionally, the emotion trajectory defines expected or normative emotional states for a user of the V/A material. Optionally, the emotion trajectory defines emotional states that are extreme and/or highly undesirable”. Para. [0032] recites “EI apparatus 102 modifies streaming of V/A material responsive to a relationship between the user's EFR generated at a time when the user is presented with the V/A frame and the FEZ for the V/A frame provided by the emotion trajectory”. Para. [0034] recites “EI apparatus 102 controls streaming of V/A frames 120 responsive to a relationship between FEZ 117 and EFR 115. Optionally, controller 107 determines a distance between EFR 115 and FEZ 117 and controls streaming of V/A material 119 responsive to the determined distance” (i.e., constructing and arranging a series of time-segmented audiovisual media content frames, or modality script, based on information from the user)); and
providing the media output to enable playback on an electronic device of the user (para. [0029] recites “In response to the user's emotional state, EI apparatus 102 modifies the progression, or streaming, of V/A material provided to user 106 by computer 105 in real time. In a situation for which the V/A material is a video game, if EI apparatus 102 receives indication that user 106 is distressed, the apparatus may modify the level of difficulty of the game, so as to provide the user with a material which will lower indications of user distress. In a situation for which the V/A material is a movie, EI apparatus 102 may modify the movie in response to the emotional state of user 106, by displaying a progression of scenes that increases the user's sense of relaxation” (i.e., providing the arranged media output to the user device)).
However, Hofshi does not explicitly teach generating at least a portion of the one or more media content items for each time-segmented directive using one or more generative artificial intelligence (Al) engines.
Labbé teaches generating at least a portion of the one or more media content items for each time-segmented directive using one or more generative artificial intelligence (Al) engines (para. [0216] recites “The MIR generator process 1900 is used to generate a MIR blueprint for an audio segment (e.g., a song) that is intended to induce a specific affective response in listeners. The MIR blueprint generated by the MIR generator process 1900 typically identifies MIR features of the song as a whole, as well as MIR features of each of multiple epochs (i.e., time-segmented features) of the audio segment, that will induce the desired affective response”. Para. [0225] recites “The functionality of the MIR generation process 1900 can also be achieved by other generative deep learning modalities like variational autoencoders (VAE) or simply a recurrent neural network (RNN) on its own. The GAN model has been evaluated as an effective means to execute the needed functionality but additional, similar algorithms may also be effective, particularly as advances in machine learning occur” (i.e., using a generating artificial intelligence engine, or model, to generate media content)).
See claim 1 for motivation to combine.
Claims 3 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Hofshi (US 20120194648 A1, herein Hofshi) in view of Labbé et al (US 20230113072 A1, herein Labbé), in further view of Dicken et al (US 20220088474 A1, herein Dicken).
Regarding claim 3, the combination of Hofshi and Labbé teaches the method of claim 1 as mentioned above.
However, the combination of Hofshi and Labbé does not explicitly teach wherein the dimensions of the ESV are each normalized to be a number from 0 to 1.
Dicken teaches wherein the dimensions of the ESV are each normalized to be a number from 0 to 1 (para. [0048] recites “data indicating player modeling in this example comprises a respective multidimensional player model 126 maintained for each player”. Para. [0069] recites “all parameters/metrics are normalized to contain values between 0 and 1”. Para. [0422] recites “the I/O components 2618 may include biometric components 2630, motion components 2634, environment components 2636, or position components 2638 among a wide array of other components” (i.e., player parameters, or dimensions, which can be associated with biometric data, can be normalized to be between 0 and 1)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine these teachings by using the method from Dicken to normalize user parameters, like the parameters used to construct the ESV from Hofshi (as modified by Labbé). Hofshi and Dicken are both directed to personalizing user experiences with media content such as games, and both Hofshi and Dicken teach that this personalization can be based on input biometric data. One of ordinary skill in the art would recognize that the user physiological parameters as taught in at least paragraph [0018] of Hofshi could be normalized using the method from Dicken.
Claim 12 is a system claim and its limitation is included in claim 3. Claim 12 is rejected for the same reasons as claim 3.
Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Hofshi (US 20120194648 A1, herein Hofshi) in view of Labbé et al (US 20230113072 A1, herein Labbé), in further view of Amores Fernandez et al* (US 20240290462 A1, herein Amores Fernandez).
*this document was included in the IDS dated 05/22/2026
Regarding claim 17, the combination of Hofshi and Labbé teaches the system of claim 10 as mentioned above.
However, the combination of Hofshi and Labbé does not explicitly teach wherein the electronic device comprises a pair of augmented-reality (AR) glasses, a virtual-reality (VR) headset, or a pair of extended-reality (XR) glasses.
Amores Fernandez teaches wherein the electronic device comprises a pair of augmented-reality (AR) glasses, a virtual-reality (VR) headset, or a pair of extended-reality (XR) glasses (para. [0055] recites “the visual output system 134 presents information using other output devices, such as a virtual or mixed reality headset 226” (i.e., a virtual reality headset)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine these teachings by modifying the user listener device from the combination of Hofshi and Labbé to utilize the virtual or mixed reality headset from Amores Fernandez. Hofshi and Amores Fernandez are both directed to generating multimedia content based on sensed emotional data from a user. As at least paragraph [0108] of Labbé states that multiple kinds of user listener devices can be supported, one of ordinary skill in the art would understand that the virtual reality headset from Amores Fernandez would fall under the broadest reasonable interpretation of a user listener device, and that the generated multimedia content from Hofshi could be presented to a user on a device such as the virtual reality headset from Amores Fernandez.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20150296228 A1 (Chen et al) teaches a method for multi-modal video segmentation to provide a user with a personalized video feed.
US 20200296480 A1 (Chappell, III et al) teaches a method for providing cinematic content comprising a directed stream of digital content based on sensed emotional data from a user.
US 20240342429 A1 (Karapantelakis et al) teaches a method for modifying aspects of an augmented or virtual reality experience based on detected sense data from a user.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LEAH M FEITL whose telephone number is (571) 272-8350. The examiner can normally be reached on M-F 0900-1700 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached on (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/L.M.F./ Examiner, Art Unit 2147
/VIKER A LAMARDO/Supervisory Patent Examiner, Art Unit 2147