Prosecution Insights
Last updated: September 17, 2026
Application No. 19/060,574

SYSTEM, METHOD AND APPARATUS FOR IMPROVING AUDIO RECORDINGS OF LIVE EVENTS

Non-Final OA §103§112
Filed
Feb 21, 2025
Priority
Feb 21, 2024 — provisional 63/556,135
Examiner
MCCORD, PAUL C
Art Unit
Tech Center
Assignee
Livewired LLC
OA Round
1 (Non-Final)
69%
Grant Probability
Favorable
1-2
OA Rounds
1y 10m
Est. Remaining
95%
With Interview

Examiner Intelligence

Grants 69% — above average
69%
Career Allowance Rate
404 granted / 584 resolved
+9.2% vs TC avg
Strong +26% interview lift
Without
With
+25.9%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
38 currently pending
Career history
621
Total Applications
across all art units

Statute-Specific Performance

§101
5.5%
-34.5% vs TC avg
§103
60.7%
+20.7% vs TC avg
§102
8.9%
-31.1% vs TC avg
§112
19.2%
-20.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 584 resolved cases

Office Action

§103 §112
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 1-20 rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. Claims 1, 11 recite a system for “receiving a plurality of livestream videos substantially simultaneously,” determining a component of each livestream, suitable to conduct processing of each of the components of each of the livestreams; determining characteristics of each component, selecting a subset of the components, presumably with respect to the corresponding livestream, mixing the subset of components, streams thereof, and “replacing the audio component of each of the livestream videos while each of the livestream videos are being streamed online.” The specification has little support for the claimed embodiment beyond restatement (please see ¶ 19 of the instant specification); and the bare assertion that such processing occurs “seamlessly,” (please see ¶ 19 of the instant specification: “videos being streamed to online viewers may be seamlessly replaced by the composite audio track,”) does little to show possession of the real or near real time, per stream, substitution of processed audio into each live stream. The specification discloses an alternate to such an embodiment a processing pipeline wherein livestream videos from plural capture devices are routed to online storage; a javascript process executes to retrieve videos, “in real, or near-real, time from the file storage system,”; a python process operates to associated two or more livestreams in accord with metadata thereof; the videos are converted in format; subject to further processing to determine audio characteristics, subset the available videos based on the characteristic; preprocess the subset of videos, presumably for mixing; combine or mix the preprocessed subset to create a composite audio track and substitute the retrieved, metadata reified, converted, processed for audio characteristics; subset reified; preprocessed, mixed and composited audio tracks are used to replace the audio in each of the subset of livestream videos while said videos are streaming (please see ¶ 67-73; Fig 5 of the instant specification). This differs substantially in scope from the claimed plurality of livestream videos and processing thereof and again does little to show possession of the real or near real time, per stream, substitution of processed audio into each live stream. Claims 2-10, 12-20 do not remedy and are similarly rejected. Claims 7, 17 additionally claim determining characteristic levels such as based on training a learning model to compare and equalize tracks in a manner not discussed in the specification particularly with regard to how such training occurs in substantially real time to add or combine tracks for delivery into each livestream. Claims 8-10, 18-20 additionally claim thresholding, component substitution, analysis and removal, steps which occur mid process and are not effectively described in the specification in the specification. Applicant is required to cancel any subject matter which lacks possession and is requested to identify any support by paragraph number and briefly describe the relevance of passages therein. Claims 1-20 rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the enablement requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to enable one skilled in the art to which it pertains, or with which it is most nearly connected, to make and/or use the invention. Claims 1, 11 require inserting a composite mix into one or more livestreams and recite “replacing the audio component of each of the livestream videos with the composite audio track while each of the livestream videos are being streamed online.” This real, near-real time substitution of audio in each outgoing stream is not found to be enabled by the specification as filed. The specification states this result twice: “a media production server receives the livestream videos in real, or near-real, time, in most cases substantially simultaneously, and processes each of the livestream videos in real or, near-real, time to generate an enhanced, composite audio track that may replace the audio component of each of the livestream videos while each of the livestream videos are being streamed to content viewers online;” (please see ¶ 19 of the instant specification) and the server “may seamlessly stop streaming the audio component of the livestream video and begin streaming the composite audio track at the fifteen second mark.” The “seamlessly” recitation serves as the entire disclosure of how the mid process splicing in of the composite stream is to be accomplished. No discussion of splicing and segment boundaries; timestamp or other synchrony management; buffering; multiplexing; or any other the necessary processes are accomplished. The only operative architecture is that discussed supra and detailed in ¶ 67-73 of the specification which discusses routing streams through memory and a plurality of distinctly provided cloud services, and leave the accomplishment of the claimed subject matter as an exercise upon the reader particular with regard to the real or near real time operations of ingesting plural streams, processing same, substituting audio therein, and outputting each processed stream while the streams are underway. This architecture is sketched out in broad strokes and no discussion of the disclosed pipeline completes in the time necessary to insert the resulting composite audio into a live stream, much less ‘each’ outgoing live stream. In fact the specification discusses the storage of the clips “after each livestream video has ended,” (please see ¶ 45 of the specification). In view of the above the claims are considered unsuitably broad. The domain of operation is latency critical and low in predictability in as much as small variations of design can be expected to produce audible discontinuities; failures of overall synchrony; buffer and jitter errors; etc.; the specification provides minimal guidance and relies heavily on named commercial services without providing details of integration, timing, etc.; thus the implementation of the claimed subject matter would rely on undue experimentation by a person having ordinary skill in the art to arrive at the splice mechanism; the synchronization issues for ‘each’ of the outgoing streams particularly in a latency compliant manner; the mid-pipeline analysis and mixing; and the entirety of the machine learning elements added in dependent claims 5, 7, 15, 17. The specification cannot be considered to provide enablement for a broadly reasonable interpretation of the claimed subject matter. Claims 2-10, 12-20 do not remedy and are similarly rejected. Claims 2, 12 additionally recite time windowed gain operations which find no enabling discussion in the specification, claims 4, 14 similarly lack disclosure of the how the weighted sum amounts to a quality score, such recitation only complicates the overall synchrony issues of the independent claims; claims 5, 7, 15, 17 similarly lack disclosure of mid-pipeline operations and how these are effected in a real or near real time manner; claims 8-10, 18-20 similarly lack clear disclosure of how the recited mid-pipeline are effected with regard to overall output as claimed. Applicant is required to cancel any subject matter which lacks enablement and is requested to identify any support by paragraph number and briefly describe the relevance of passages therein. The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1-20 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention. Claims 1, 11 recite “each of the plurality of audio components,” in an ambiguous manner as it cannot be rightly discerned whether the plurality refers to the recited “each” of the audio components of each livestream video either inclusive or exclusive of the recited “an audio component;” or if “each,” of the items exists in the selected “subset of the audio components,” etc., and as such the claims are considered indefinite. Claims 2-10, 12-20 do not remedy and are similarly indefinite. In addition, claims 2, 9, 12, 19 recite “the selected audio components of the subset of audio components,” in such a way as it may be inferred that these are audio components common within the recited selected “a subset of the audio components,” of claims 1, 11—such an interpretation does not improve the clarity of the relevant claims. Similarly claim 3, 8, 10, 13, 18, 20 ambiguously conflate the terms “audio quality threshold,” “audio characteristic threshold,” and “predetermined quality threshold,”—as the specification discusses various thresholds, consistent terminology with resolving clear antecedents thereof would improve the clarity of the relevant claims; further the first and second thresholds of claims 3, 13 resolve groupings of audio characteristics in an unclear manner. Claims 7, 17 discuss a “sonic tone,” in a manner which resolves no antecedent in the specification and thus it must be surmised what a sonic tone comprises—does Applicant intend the claim to produce a result such as by a match of equalization values to a reference, to a characteristic tone of a genre; such as to emulate a transfer function of the reference song; such as to …, etc.? Claims 8, 18 appear to require a second composite audio track, as the recited mixing of the claim occurs during provision of the recited “composite audio track,” such as recited in claim 1 the disposition, availability, etc. of the composite audio track for subsequent mixing is unclear. Claim 9, 10, 19, 20 again resolve the recited audio components unclearly. Appropriate correction is required. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1-20 rejected under 35 U.S.C. 103 as being unpatentable over Ojanpera: 20130297054 hereinafter Oj further in view of Dean: 20200234733 hereinafter De and further in view of Rider: 20160191591 hereinafter Ri. Regarding claim 1 Oj teaches: A method, performed by an online media production server, for improving audio quality of livestream videos sourced from live events (Oj: Abstract; ¶ 1-3, 76, 92-97, 176; Fig 2: system comprises a processor operable to execute instruction stored in and instantiated from memory and operable to improve rendering of user generated content of a live event including video of a live event), comprising: receiving a plurality of livestreams substantially simultaneously from a plurality of spectators at an event, each of the livestreams comprising an audio component (Oj: Abstract; ¶ 1-6, 76-79, 152; Fig 1, 2, 12: plurality of user devices arrayed about an event operate to record a same event at a same time from a plurality of positions and upload individual user recording such as using a network interface of the user device(s) to transmit rather than locally store the event audio, video, etc. said transmission occurring at substantially similar, simultaneous, etc. time frames and for at least the purpose of allowing access to a timeline of overlapping sound sources from which to stitch, splice, etc. together or otherwise assemble a timewise consistent media); evaluating each of the audio components of each livestream to identify one or more audio characteristics of each audio component (Oj: ¶ 7, 80, 111-119, 142, 162-163; Fig 1, 2, 12: such as by server-side analysis of each receive audio signal to determine characteristics thereof such as by analyzing each/any of the sound sources to determine position, location, orientation, dominance, focus, etc. data thereof); selecting a subset of the audio components for mixing based on the one or more audio characteristics of each of the plurality of audio components (Oj: ¶ 6, 169-172;; Fig 1, 2, 12-17: a downmix source selector receives audio source data, metadata thereof, determines a suitable number of appropriate sources for a high quality downmix such as by analyzing sound parameters to determine audio signals dependent from a dominant signal source with respect to a particular listening position, location, etc. parameters); mixing the subset of the audio components to produce a composite audio track (Oj: ¶ 86, 108; Fig 1, 2, 12: appropriate signal sources passed to a downmixer configured to use the selected appropriate audio sources to generate a composite, mixed, etc. signal for transmission to downstream user devices). Oj strongly suggests but does not explicitly teach the upload of livestream videos nor the replacing the audio component of each of the livestream videos with the composite audio track while each of the livestream videos are being streamed online. In a related field of endeavor De teaches a system and method comprising a processor operable for synchronizing a plurality of low quality audio content by replacing portions thereof with higher quality audio portions (De: Abstract: ¶ 1, 2, 8, 30, 79; Figs 1-3: system comprises a processor operable to execute instruction stored in and instantiated from memory and operable to improve rendering of live event content by replacing portions of real time uploaded, live streamed event audio with an improved version thereof, synchronizing same) comprising receiving a plurality of livestream videos substantially simultaneously from spectators at an event, each comprising an audio component (De: ¶ 2, 8, 12, 30-34, 37, 48: multiple users at a same event operate respective user devices to record media including audio and video of the event, upload, stream, livestream etc. same such as in concert with time, geo location etc. metadata identifying the event of performances and the part of the performance recorded) said system operative for evaluating the audio components of each live stream (De: ¶ 29, 31, 41: each livestream audio identified; assessed for quality; matched to the source content such as using fingerprint data thereof; etc.); and mixitively replacing the audio component of each of the livestream videos with the composite audio track while each of the livestream videos are being streamed online (De: ¶ 29-36, 45, 56; Fig 1-3, 8: system operates to replace portions of a user recorded audio with audio from a better quality audio source, stream same). It would have been obvious to one of ordinary skill in the art before the effective filing date of the instant application to modify the audio source analysis, subset selection and downmixing of Oj to thereby improve content in the manner taught or suggested by the De system and method for at least the purpose of generating a plurality of livestream videos with improved audio quality where possible; one of ordinary skill in the art would have expected only predictable results therefrom. It may be that Oj in view of De does not explicitly teach the recited simultaneous aspects of the recited receiving and replacing of audio for each livestream. In a related field of endeavor Ri teaches a system and method for live crowdsourced media streaming (Ri: Abstract; ¶ 13, 40, etc.) comprising; receiving a plurality of livestream videos substantially simultaneously from a plurality of spectators at an event; each comprising an audio component (Ri: ¶ 1, 13, 23, 45-55: potentially thousands of spectator devices operable to capture video and audio streams of an event, generate a live stream thereof such as by streaming contemporaneously with capture to a computing device for presentation thereon), evaluating audio components of each stream to identify and generate characteristics thereof (Ri: ¶ 44: streams analyzed to determine metadata thereof such as image quality data, relevance to the event, view angle, etc.); selecting a subset of components to produce a composite video (Ri: ¶ 44, 45: event match or relevance used to reify the subset of streams and generate an enhanced stream of the event); and managing the generation of per stream improvements for each of the livestream videos while each are being streamed online (Ri: ¶ 29, 40, 45, 50; Fig 3: system mixitive combines the media streams uploaded from the user devices by forking the input and output to downstream presentation devices to stream enhanced video substantially contemporaneously with their receipt). It would have been obvious to one of ordinary skill in the art before the effective filing date of the instant application to incorporate the Oj in view of De system and method into the near real time media streaming platform of Ri for at least the purpose of modifying audio of the input and output streams in a timewise parsimonious manner to thereby enhance a downstream users exercise of the livestreamed media; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 2 Oj in view of De in view of Ri teaches or suggests: The method of claim 1, wherein mixing the subset of the audio components comprises: aligning each of the selected audio components in time (Oj: ¶ 152, 161: system maps plurality of videos upon a common timeline); (De: ¶ 37: system synchronizes independently recorded content); (Ri: ¶ 50: such as based on temporal correspondence of the plural livestreams) increasing a gain of a first audio component of the subset of audio components during a time when a first audio characteristic of the first audio component during the time exceeds a second audio characteristic of a second audio component of the subset of audio components during the time (Oj: ¶ 2, 5, 124, 142: system mixings in dominant sounds such as to reduce background noise, improve signal quality, impart a desired gain profile, etc.); and adding the gain-increased first audio component and the second audio component together to produce the composite audio track (Oj: ¶ 5, 108: downmixer combines selected sources into a composite audio stream); (De: ¶ 31, 32, 46, etc.: system operates to combine source and remote audio, said audio comprising differing gain profiles). Oj in view of De in view of Ri does not explicit teach increasing a gain of a specific audio component at particular times based on the relationship of the specific audio component to the audio components of other streams however Examiner takes official notice that mixing was well known in the art before the effective filing date of the invention to comprise exactly the recited adjustments of volume of one media stream based on relationships of the parameters thereof with the parameters of other media streams. The claim is thus considered obvious over Oj as modified by De, and Ri as addressed in the base claim as it would have been obvious to apply the further teaching of Oj, De, and/or Ri to the modified device of Oj, De and Ri; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 3 Oj in view of De in view of Ri teaches or suggests: The method of claim 1, wherein selecting the subset of the audio components comprises: defining one or more audio characteristic thresholds each associated with one of the one or more audio characteristics (Oj: ¶ 132-135, 164: such as selection of a particular stream for mixing, combining, etc. based on thresholded location of audio characteristics); determining, for each of the plurality of audio components, one or more audio characteristic levels (Oj: ¶ 116-119, 142, etc. such as determining of signal energy, direction, dominance, etc.); (De: ¶ 29: such as determining quality levels of audio); (Ri: ¶ 44: such as determining quality of a stream); selecting a first audio component for mixing when a first audio characteristic level exceeds a first audio characteristic threshold; and selecting a second audio component for mixing when a second audio characteristic level of a second audio characteristic exceeds a second audio characteristic threshold (Oj: ¶ 164-172; system selects sources based on thresholded location to manage focus determination). The claim is considered obvious over Oj as modified by De, and Ri as addressed in the base claim as it would have been obvious to apply the further teaching of Oj, De, and/or Ri to the modified device of Oj, De and Ri; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 4 Oj in view of De in view of Ri teaches or suggests: The method of claim 1, wherein selecting the subset of the audio components comprises: calculating, for each of the plurality of audio components, an audio score comprising a weighted sum of one or more of the audio characteristic levels (Oj: ¶ 142, 163-165: system scores characteristic of stream which drives selection thereof such as by determining a dominant event and focus to thereby aggregate directional and energy characteristics of streams); (Ri: ¶ 44-50: system maintains a per source stream score and selects streams based thereon, such as for matching by image quality); selecting a first audio component for mixing when a first audio score of one of the plurality of audio components is greater than audio scores of any other audio component; and selecting a second audio component for mixing when a second audio score of another of the plurality of audio components is greater than audio scores of each of the remaining audio components (Oj: ¶ 169-172: system selects highest ranking streams of a plurality of streams); (Ri: ¶ 47: system searches for streams based on ranks, scores, etc. thereof). The claim is considered obvious over Oj as modified by De, and Ri as addressed in the base claim as it would have been obvious to apply the further teaching of Oj, De, and/or Ri to the modified device of Oj, De and Ri; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 5 Oj in view of De in view of Ri teaches or suggests: The method of claim 1, further comprising: using a machine learning, algorithmic, etc. model (Oj: ¶ 7, 111, 131-135, 159-174, etc.: system uses a classification processes, itself an algorithmic, machine learning type process or trained model therefor); (De: ¶ 29-36, 40-45, 56; Fig 1-3, 8: system operates to replace portions of a user recorded audio with audio from a better quality audio source, stream same; such as based on matching of an audio source stream to an reference stream based on audio fingerprint of the source stream); (Ri: ¶ 48: system utilizes a machine learning model) to: determine one or more audio characteristic levels of the plurality of audio components (Oj: ¶ 7, 111, 131-135, 159-174, etc.: system determines and selects streams based on characteristic levels thereof); and select one or more of the plurality of audio components for mixing based on the one or more audio characteristic levels (Oj: ¶ 6, 148, 163-165: system replaces a selected source when a better source meets a selection criterion; said better source selected from the large number of livestreamed media); (De: ¶ 29-36, 40-45, 56; Fig 1-3, 8: system operates to replace portions of a user recorded audio with audio from a better quality audio source, stream same; such as based on determined matching audio to an audio fingerprint of a stream); (Ri: ¶ 45-47: system replaces delivered stream based on evaluations thereof). While Oj in view of De in view of Ri does not explicitly teach training a machine learning model to perform the recited algorithmic processes Examiner takes official notice that training of machine learning models for classification, matching, volume and frequency adjustment, mixing, etc. based on identifying particular stream of music, songs, etc.; such as using parameters; fingerprints or other metadata; acoustical properties; etc. thereof were well known in the art before the effective filing date of the instant invention and would have comprised an obvious inclusion for the purpose of implementation of the taught algorithmic methods thereby. The claim is thus considered obvious over Oj as modified by De, and Ri as addressed in the base claim as it would have been obvious to apply the further teaching of Oj, De, and/or Ri to the modified device of Oj, De and Ri; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 6 Oj in view of De in view of Ri teaches or suggests: The method of claim 1, further comprising: identifying a song from one or more of the plurality of audio components (Oj: ¶ 7, 111, 131-135, 159-174, etc.: system uses a classification processes, itself an algorithmic, machine learning type process or trained model therefor); (De: ¶ 29-36, 40-45, 56; Fig 1-3, 8: system operates to replace portions of a user recorded audio with audio from a better quality audio source, stream same; such as based on determined matching of an audio source stream to an reference stream based on audio fingerprint of the source stream); (Ri: ¶ 48: system utilizes a machine learning model); comparing the composite audio track to a reference song associated with the song ; and equalizing the composite audio track to match amplitudes and frequencies of the reference song (De: ¶ 32, 41-45: system identifies an audio source stream by matching of the audio source stream to an reference stream based on audio fingerprint of the source stream and performs filtering or other typical adjustments using sound engineering equipment therefor); (Ri: ¶ 43-47: system analyzes media stream to identify metadata parameters thereof and extract searchable therefrom and identifies matching media based thereon). The claim is considered obvious over Oj as modified by De, and Ri as addressed in the base claim as it would have been obvious to apply the further teaching of Oj, De, and/or Ri to the modified device of Oj, De and Ri; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 7 Oj in view of De in view of Ri teaches or suggests: The method of claim 1, further comprising: using a machine learning, algorithmic, etc. model(Oj: ¶ 7, 111, 131-135, 159-174, etc.: system uses a classification processes, itself an algorithmic, machine learning type process or trained model therefor); (De: ¶ 29-36, 40-45, 56; Fig 1-3, 8: system operates to replace portions of a user recorded audio with audio from a better quality audio source, stream same; such as based on matching of an audio source stream to an reference stream based on audio fingerprint of the source stream); (Ri: ¶ 48: system utilizes a machine learning model) to: identify a song from one or more of the plurality of audio components to: compare the composite audio track to a reference song associated with the song (De: ¶ 41-45: system identifies a source stream based on matching of an audio source stream to an reference stream based on audio fingerprint of the source stream such as by audio or acoustic fingerprint comparisons of events in an event database); (Ri: ¶ 43-47: system analyzes media stream to identify metadata parameters thereof and extract searchable therefrom and identifies matching media based thereon); and equalize the composite audio track to match a sonic tone of the reference song (De: ¶ 32: such as by performing filtering or other typical adjustments using sound engineering equipment therefor). While Oj in view of De in view of Ri does not explicitly teach training a machine learning model to perform the recited algorithmic processes Examiner takes official notice that training of machine learning models for classification, matching, volume and frequency adjustment, mixing, etc. based on identifying particular stream of music, songs, etc.; such as using parameters; fingerprints or other metadata; acoustical properties; etc. thereof were well known in the art before the effective filing date of the instant invention and would have comprised an obvious inclusion for the purpose of implementation of the taught algorithmic methods thereby. The claim is considered obvious over Oj as modified by De, and Ri as addressed in the base claim as it would have been obvious to apply the further teaching of Oj, De, and/or Ri to the modified device of Oj, De and Ri; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 8 Oj in view of De in view of Ri teaches or suggests: The method of claim 1, further comprising: receiving a first livestream video while the composite audio track is being provided online, the first livestream video comprising a first audio component (Oj: Abstract; ¶ 1-6, 76-79, 113-116, 152; Fig 1, 2, 12: plurality of users transmit livestream audio such as by a population changing over the course of an event); (De: ¶ 2, 8, 12, 30-34, 37, 40-43, 48: multiple users livestream event media which is reified timewise for placement on a timeline; each stream additionally used for determining metadata including acoustic metadata such as fingerprints for matching with reference fingerprints); (Ri: ¶ 6-7, 45, 50: system receives plurality of user streams over the course of an event); analyzing the first audio component to determine an audio characteristic of the first audio component (Oj: ¶ 113-116: system analyzes each/any incoming stream); (De: ¶ 41, etc.: system determines quality, fingerprints, etc. of plurality of audio streams); (Ri: ¶ 44: system analyzes each incoming stream); determining that the audio characteristic exceeds an audio quality threshold (Oj: ¶ 6, 131-135, 164: system performs threshold comparison of analyzed characteristic such as to determine most relevant sound sources); (Ri: ¶ 47-49: such as by ranking or scoring input media by user to select media) and mixing the first audio component with the subset of audio components to produce the composite audio track (Oj: 5, 6, 108, etc.: system includes qualifying source into a particular downmix); (De: ¶ 31, etc.: mixer combines determined sources); (Ri: ¶ 45: newly arrived media dynamically included into delivered output). The claim is considered obvious over Oj as modified by De, and Ri as addressed in the base claim as it would have been obvious to apply the further teaching of Oj, De, and/or Ri to the modified device of Oj, De and Ri; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 9 Oj in view of De in view of Ri teaches or suggests: The method of claim 1, further comprising: receiving a first livestream video while the composite audio track is being provided online, the first livestream video comprising a first audio component (Oj: Abstract; ¶ 1-6, 76-79, 113-116, 152; Fig 1, 2, 12: plurality of users transmit livestream audio such as by a population changing over the course of an event); (De: ¶ 2, 8, 12, 30-34, 37, 48: multiple users livestream event media which is reified timewise for placement on a timeline); (Ri: ¶ 6-7, 45, 50: system receives plurality of user streams over the course of an event); analyzing the first audio component to determine an audio characteristic of the first audio component (Oj: ¶ 113-116: system analyzes each/any incoming stream); (De: ¶ 41, etc.: system determines quality of plurality of audio streams); (Ri: ¶ 44: system analyzes each incoming stream)and substituting one of the selected audio components of the subset of audio components with the first audio component during mixing when the audio -characteristic of the first audio component exceeds an audio characteristic of one of the subset of audio components (Oj: ¶ 6, 148, 163-165: system replaces a selected source when a better source meets a selection criterion; said better source selected from the large number of livestreamed media); (De: ¶ 29-36, 45, 56; Fig 1-3, 8: system operates to replace portions of a user recorded audio with audio from a better quality audio source, stream same); (Ri: ¶ 45-47: system replaces delivered stream based on evaluations thereof). The claim is considered obvious over Oj as modified by De, and Ri as addressed in the base claim as it would have been obvious to apply the further teaching of Oj, De, and/or Ri to the modified device of Oj, De and Ri; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 10 Oj in view of De in view of Ri teaches or suggests: The method of claim 1, further comprising: as the livestream videos with the composite audio track are being provided online (Oj: Abstract; ¶ 1-6, 76-79, 113-116, 152; Fig 1, 2, 12: plurality of users transmit livestream audio such as by a population changing over the course of an event); (De: ¶ 2, 8, 12, 30-34, 37, 48: multiple users livestream event media which is reified timewise for placement on a timeline); (Ri: ¶ 6-7, 45, 50: system receives plurality of user streams over the course of an event), continuing to evaluate each of the plurality of audio components (Oj: Abstract; ¶ 1-6, 76-79, 113-116, 148, 152; Fig 1, 2, 12: etc.: system performs ongoing analysis of changing source population of media); (Ri: ¶ 40-45: system receives and analyzes media concurrent with the delivery thereof) to identify when an audio characteristic of any of the plurality of audio components degrades past a predetermined quality threshold (Oj: ¶ 6, 131-135, 164-174: system performs ongoing analysis of plurality of streams, threshold comparison of analyzed characteristics thereof such as to determine most relevant sound sources including selection of based on quality wherein relationship of a particular stream to a focus threshold, i.e. in or out of focus, comprises a quality determination and drives selection of streams); (Ri: ¶ 40-45: system merges or eliminates particular streams based on quality thereof); identifying a first audio component comprising an audio characteristic that has degraded past the predetermined quality threshold (Oj: ¶ 6, 131-135, 159, 164-174: such as by identifying in and out of focus sources); (Ri: ¶ 40-45); and removing the first audio component from being mixed with the subset of audio components when the audio characteristic of the first audio component has degraded past the predetermined quality threshold (Oj: ¶ 6, 131-135, 164-174: such as by removing out of focus sources, replacing same with in focus sources); (De: ¶ 29-36, 45, 56; Fig 1-3, 8: system operates to replace portions of a user recorded audio with audio from a better quality audio source, stream same) ; (Ri: ¶ 45-47: system replaces delivered stream based on evaluations thereof). The claim is considered obvious over Oj as modified by De, and Ri as addressed in the base claim as it would have been obvious to apply the further teaching of Oj, De, and/or Ri to the modified device of Oj, De and Ri; one of ordinary skill in the art would have expected only predictable results therefrom. Regarding claim 11—the claim is considered to recite substantially similar subject matter to that of claim 1 and is similarly rejected. Regarding claim 12—the claim is considered to recite substantially similar subject matter to that of claim 2 and is similarly rejected. Regarding claim 13—the claim is considered to recite substantially similar subject matter to that of claim 3 and is similarly rejected. Regarding claim 14—the claim is considered to recite substantially similar subject matter to that of claim 4 and is similarly rejected. Regarding claim 15—the claim is considered to recite substantially similar subject matter to that of claim 5 and is similarly rejected. Regarding claim 16—the claim is considered to recite substantially similar subject matter to that of claim 6 and is similarly rejected. Regarding claim 17—the claim is considered to recite substantially similar subject matter to that of claim 7 and is similarly rejected. Regarding claim 18—the claim is considered to recite substantially similar subject matter to that of claim 8 and is similarly rejected. Regarding claim 19—the claim is considered to recite substantially similar subject matter to that of claim 9 and is similarly rejected. Regarding claim 20—the claim is considered to recite substantially similar subject matter to that of claim 10 and is similarly rejected. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to PAUL C MCCORD whose telephone number is (571)270-3701. The examiner can normally be reached 730-630 M-F. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, CAROLYN EDWARDS can be reached at (571) 270-7136. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PAUL C MCCORD/ Primary Examiner, Art Unit 2692
Read full office action

Prosecution Timeline

Feb 21, 2025
Application Filed
Sep 02, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12724972
AUTOMATIC SENTENCE CONDITION MATCHING USING NATURAL LANGUAGE PROCESSING
2y 11m to grant Granted Sep 01, 2026
Patent 12718824
Audio Signal Encoding Method, Decoding Method, Encoding Device, and Decoding Device
3y 10m to grant Granted Aug 25, 2026
Patent 12718804
DOMAIN SPECIALTY INSTRUCTION GENERATION FOR TEXT ANALYSIS TASKS
3y 1m to grant Granted Aug 25, 2026
Patent 12718789
Inference-time Control of Transformers for Audio Generation
2y 5m to grant Granted Aug 25, 2026
Patent 12710919
Multiple Groupings in a Playback System
2y 1m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
69%
Grant Probability
95%
With Interview (+25.9%)
3y 5m (~1y 10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 584 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month