Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
CLAIM INTERPRETATION
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are:
“an audio input interface configured to…” “the audio metadata manager configured to…” “an audio analysis module configured to…” “a generative AI model configured to…” “a storytelling algorithm configured to…” “a graphics processing unit configured to…” “the visual element library configured to…” in Claim 16.
Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-6 and 13-19 are rejected under 35 U.S.C. 103 as being unpatentable over Bendale et al. (US2021/0201549 A1) in view of McNulty et al. (US 2024/0361827 A1).
As to Claim 1, Bendale teaches A method for presenting an artificial intelligence (AI)-enhanced visual experience (Bendale, Abstract), comprising:
receiving, by an audio input interface, audio content from an audio content source (Bendale discloses “In particular embodiments, the one or more computing systems may receive one or more inputs. In particular embodiments, the one or more inputs may include (but is not limited to) one or more non-video inputs. The non-video inputs may comprise at least one of a text input, an audio input, or an expression input” in [0029]; microphone in [0067]);
processing, by an audio analysis module, the audio content to obtain audio analysis outputs (Bendale discloses “In particular embodiments, analysis may be performed on text input, audio input, expression input, and video inputs to identify characteristics of a particular semantic context” in [0030]);
translating, by a generative Al model, a first portion of the audio analysis outputs to obtain an audio semantic description (Bendale discloses “Certain inputs may be grouped together when training a machine-learning model as described herein to form a particular semantic context. In particular embodiments, each of the plurality of semantic contexts may be indicative of an expression” in [0030]);
mapping, by a video conversion engine, at least the audio semantic description and the storytelling theme to a visual element set; selecting, by the video conversion engine, one of a selection group comprising the visual element set and a visual element subset of the visual element set (Bendale discloses “In particular embodiments, each of the plurality of semantic contexts may be indicative of an expression. As an example and not by way of limitation, one of the plurality of semantic contexts may include a sad expression.” in [0030]; “As an example and not by way of limitation, for a user input "Hello, how are you?" the one or more computing systems may identify nodes corresponding to several different semantic contexts, such as a nodding semantic context, a talking semantic context, and a smiling semantic context” in [0031]; “In particular embodiments, each node 1116 may correspond to a particular semantic context and include one or more expressions, behaviors, actions associated with the node 1116. The expressions, behaviors, and actions associated with each node 1116 may be used to generate an output video 1118 of a digital human/avatar.” in [0056].);
creating, by a storytelling algorithm, a visual narrative using the one of the selection group, the visual narrative comprising a video constructed by the Al; and displaying, by a visual controller, the visual narrative through a visual content target (Bendale discloses “The text output and/or audio output of the answer 1006 may be processed with a graph of the digital human system 200 to identify a semantic context and associated actions to be performed by a digital human/avatar. The graph may be used to identify one or more nodes that correspond to the answer 1006. The video output of the digital human 804 may be sent to the end user 202 through both a visual and audio feed.” in [0055].)
Bendale doesn’t directly use claim language “storytelling theme”. The combination of McNulty further teaches following limitations:
processing, by the generative Al model, a second portion of the audio analysis outputs and the audio semantic description to create a storytelling theme (Bendale discloses “The text output and/or audio output of the answer 1006 may be processed with a graph of the digital human system 200 to identify a semantic context and associated actions to be performed by a digital human/avatar. The graph may be used to identify one or more nodes that correspond to the answer 1006. The video output of the digital human 804 may be sent to the end user 202 through both a visual and audio feed” in [0055]. McNulty further discloses “In the context of music, to curate means to carefully select, organize, and present musical works or performances in a thoughtful and coherent manner. This can involve creating playlists, assembling a lineup for a concert or festival, or even designing a music program for a specific event. The goal of music curation is to provide a meaningful and engaging experience for the audience, showcasing the artistic vision of the curator and often highlighting specific themes, genres, or styles” in [0173]; “Themed sound walks or tours” in [1327] and “Audio storytelling and narrative experiences” in [1328].)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the invention of Bendale with the teaching of McNulty so as to create immersive audio experiences to enhance storytelling (McNulty, [1328]).
As to Claim 2, Bendale in view of McNulty teaches The method of Claim 1, wherein the audio analysis outputs comprise audio waveform components, audio complex attributes, audio embedded contexts, and audio emotional undertones (Bendale discloses “In particular embodiments, analysis may be performed on text input, audio input, expression input, and video inputs to identify characteristics of a particular semantic context. In particular embodiments, each of the plurality of semantic contexts may be indicative of an expression…” in [0030]. McNulty further discloses waveform components in [1331]; complex attributes in [0428]; context in [0417] and emotion expression in [0513].)
As to Claim 3, Bendale in view of McNulty teaches The method of Claim 2, wherein the first portion of the audio analysis outputs comprises one or more of: the audio waveform components and the audio complex attributes (McNulty, [0428, 1331].)
As to Claim 4, Bendale in view of McNulty teaches The method of Claim 2, wherein the second portion of the audio analysis outputs comprises one or more of: the audio embedded contexts and the audio emotional undertones (Bendale, [0027, 0030]. McNulty, [0096, 0110, 0417].)
As to Claim 5, Bendale in view of McNulty teaches The method of Claim 1, wherein selecting the one of the selection group is based on one or more of: an audio mood, audio source information comprising an audio genre and an audio rhythm, and a recurring theme (McNulty discloses “10. Sound: Curating sound can involve selecting and arranging different sound sources or adjusting volume and tone to create a particular mood or ambiance. Curating sound can also involve manipulating sound to create rhythm, harmony, or dissonance.” in [0070], see also [0024-0025].)
As to Claim 6, Bendale in view of McNulty teaches The method of Claim 5, wherein processing the audio content further obtains one or more of: the audio mood, an audio intensity, an audio expressiveness, and the recurring theme (McNulty discloses “Voice tone analysis: Microphones can capture voice data, and software algorithms can analyze changes in pitch, intensity, and speech rate to infer emotions” in [2283]; “Advanced voice cloning systems are trained to understand and reproduce these emotional cues, ensuring that the cloned voice conveys a similar emotional depth and expressiveness as the original speaker's voice” in [0513], see also [0025].)
As to Claim 13, Bendale in view of McNulty teaches The method of Claim 1, wherein the visual content target comprises one of: a display screen, a media projector, a virtual reality device, and an augmented reality device (Bendale discloses “The audio-visual content generated using the disclosed technology may be rendered with wide ranging end-point devices, such as smartphones, wearable devices, TVs, digital screens, holographic displays, or any other media consumption device” in [0025].)
As to Claim 14, Bendale in view of McNulty teaches The method of Claim 1, further comprising, after displaying the visual narrative: receiving, by an interactive feedback system, ambient feedback from an ambient feedback source; and adjusting, by the storytelling algorithm, the visual narrative based on the ambient feedback (Bendale discloses “Herein disclosed are one or more approaches for creating lifelike digital humans that may have sensing, interaction, understanding, and cognition capabilities like real humans, while at the same time being reactive, controllable, and having varying degrees of autonomous behavior for decision making” in [0026]. McNulty discloses “Curating temperature can also involve manipulating temperature to create comfort, energy efficiency, or safety” in [0071]; “Sentiment analysis touch sensor: A sensor embedded in everyday objects that can detect and analyze the emotional state of a person through their touch, such as the pressure, temperature, and patterns of contact” in [2292]; “multiple variable biometric measurements are detected live and real time, an instant and continuous reaction by AI may be generated” in [2807], see also [2284-2289].)
As to Claim 15, Bendale in view of McNulty teaches The method of Claim 1, further comprising, after displaying the visual narrative: receiving, by a customization and control interface, a manual visual change from a user; and adjusting, by the storytelling algorithm, the visual narrative based on the manual visual change (McNulty discloses “Customization options: Allowing users to manually adjust and customize their content preferences can provide valuable information to the AI model, helping it understand the user's unique requirements” in [2265].)
Claim 16 recites similar limitations as claims 1 & 14 but in a device form. Therefore, the same rationale used for claims 1 $&14 is applied.
As to Claim 17, Bendale in view of McNulty teaches The live entertainment platform of Claim 16, wherein: the audio content originates from an audio content source, and the audio content source comprises one of: a microphone, disc-jockey (DJ) equipment, a musical instrument, and a digital music file (McNulty, [0174].)
Claim 18 is rejected based upon similar rationale as Claim 13.
As to Claim 19, Bendale in view of McNulty teaches The live entertainment platform of Claim 16, wherein: the ambient feedback originates from an ambient feedback source, and the ambient feedback source comprises one of: a microphone, a camera, social media, and an environmental sensor (Bendale discloses microphone, camera in [0067] and environment sensor in [0043]. McNulty discloses social media in [0011],)
Claims 7-9 are rejected under 35 U.S.C. 103 as being unpatentable over Bendale in view of McNulty and Yip et al. (US 2025/0177864 A1).
As to Claim 7, Bendale in view of McNulty teaches The method of Claim 5, further comprising:
determining, by the audio input interface and concurrent with or after reception of the audio content, audio content metadata from the audio content source; and parsing, by the audio analysis module, the audio content metadata to obtain one or more of: the audio source information, audio temporal data, audio technical specifications, and audio contextual information (Bendale discloses semantic context in [0030]; “Graph Q&R 262 may query graphs 234 of the foundry 208 in the form of metadata and/or keypoint queries and receives audio visual data from the graphs 234” in [0042]. Here, metadata can include audio metadata. For example, Yip discloses “In some implementations, the additional characteristics can be extracted from metadata embedded with the audio signal in the audio data… For example, in the case of music audio, metadata included with the audio signal can be used to identify characteristics, such as audio signal identifier, tone, beat, lyrics, speed/pace, genre, title, artist, composer, track number, popularity index, etc.” in [0040]; “a temporal characteristic of the audio signal and descriptive characteristics associated with the audio data” in [0018], see also [0038].)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the invention of Bendale and McNulty with the teaching of Yip so as to analyze the audio content and associated metadata to determine the context, tone, genre, lyrics etc. (Yip, [0038]).
As to Claim 8, Bendale in view of McNulty and Yip teaches The method of Claim 7, further comprising, after parsing the audio content metadata, processing, by the audio analysis module and to produce an enhanced audio understanding, one or more of: the audio source information, the audio temporal data, the audio technical specifications, and the audio contextual information (McNulty discloses “Collecting and analyzing biometric data at different time points allows models to better understand and adapt to changes in a person's preferences, emotional states, and needs” in [2403], see also [2408]. Yip, [0018, 0038, 0054].
As to Claim 9, Bendale in view of McNulty and Yip teaches The method of Claim 7, further comprising, prior to displaying the visual narrative, aligning, by the storytelling algorithm, the visual narrative with the audio content based on the audio temporal data (McNulty discloses verification data such as timestamps in [1916]. Yip further discloses “The select ones of the characteristics of the audio data include descriptive data that provides details of the audio data and at least one temporal data that can be used to match the audio data to corresponding game scene.” in [0029].)
Claims 10-12 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Bendale in view of McNulty and Carlsson et al. (US 2015/0215691 A1).
As to Claim 10, Bendale in view of McNulty teaches The method of Claim 2, further comprising, concurrent with displaying the visual narrative, displaying, by a second visual controller and through a second visual content target, a lighting sequence based on one or more of: the audio waveform components and audio sound types (Bendale discloses “In particular embodiments, the keypoints may be landmark points or interest points detected by various computer vision and video analysis processes that highlight and characterize stable and reoccurring regions in video. The landmark points may be called keypoints and are tracked throughout a video” in [0040]; “In particular embodiments, the audio to KP transformer 280 may be a machine learning model that predicts keypoints given a set of incoming audio stream. In particular embodiments, the keypoints may undergo further transformation such as mixing and animating keypoints from a predefined set of animation curves and statistical models” in [0045]. Carlsson further discloses “…to create a lighting experience that accompanies music playback on the system. Lighting sequences may be created based on music genre, beat, etc.” in [0003].)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the invention of Bendale and McNulty with the teaching of Carlsson so that the individual speakers of the network have lamps on them that are controlled to present a light show in synchrony with the audio being played by the system (Carlsson, Abstract).
As to Claim 11, Bendale in view of McNulty and Carlsson teaches The method of Claim 10, wherein processing the audio content further obtains the audio sound types (Carlsson, [0044] and Fig 4.)
As to Claim 12, Bendale in view of McNulty and Carlsson teaches The method of Claim 10, wherein the second visual content target comprises a lighting device (Carlsson discloses lamps are controlled to present a light show in Abstract.)
Claim 20 is rejected based upon similar rationale as Claims 10 & 12.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WEIMING HE whose telephone number is (571)270-1221. The examiner can normally be reached on Monday-Friday, 8:30am-5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tammy Goddard can be reached on 571-272-7773. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WEIMING HE/
Primary Examiner, Art Unit 2611