Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-2, 5, 11-13, and 16 are rejected under 35 U.S.C. 102(a)(1) as being anticipated over Pramod et al. (WO 2024145444 A1 ).
Regarding claim 1, Pramod teaches:
An audio device, comprising (see [0022]:” computing system 100 is incorporated into an
audio device, such as an audio player, audio/video player, media player, smart phone, tablet, laptop computer, desktop computer, an in-vehicle system, and/or the like”):
at least one processor configured to (see [0022]:” Computing device 102 includes, without limitation, an I/O interface 106, a processor 108…”, also see Element 108, Fig. 1):
receive an audio input signal (the signal from audio input is an audio input signal, see [0025]:”…audio input can be received through the I/O interface 106…”);
extract content-based features from the audio input signal (audio content type identification corresponds to content-based features, and is in audio frames/segmentations which is derived from the audio input data, see “Audio scene analysis based on audio content type identification …segmenting the audio data into a plurality of audio frames; extracting one or more audio features from the plurality of audio frames…”, Abstract; module 202 performs feature extraction, i.e. extracting content-based features, see [0036]:“…feature extraction module 202 uses signal processing techniques to extract features, such as the frequency spectrum and/or the like, which provides insight into the pitch and harmonic content of the segments of audio data 116…”, also see Fig. 2);
provide the extracted content-based features to a machine learning (ML) based content detection module running at the audio device (CNN is a machine learning model processing extracted content based features, therefore it’s a machine learning based content detection module. CNN is running on the on-device audio scene analyzer, therefore running at the audio device, see [0042-0043]: ”FIG. 3 is a block diagram of audio type classifier 210 included in audio scene analyzer 114 processing extracted features 204, according to various embodiments. As illustrated, audio type classifier 210 includes, without limitation, a convolutional neural network (CNN) 300… CNN 300 analyzes extracted features 204 for an audio segment…”, also see Fig. 1);
adjust an audio output setting for the output based on a content indicator from the ML based content detection module (confidence score is at least part of a content indicator, see [0032]: “Audio settings application 118 receives the audio scene type determined by audio scene analyzer 114 and the confidence score corresponding to the audio scene type. Audio settings application 118 then uses the determined audio scene type and the confidence score to select or adjusts audio parameters, such as equalizer settings, volume levels, and acoustic effects, tailored to enhance audio data 116 for the current audio segment”).
Regarding claim 2, Pramod teaches the ML based content detection module is configured to provide the content indicator independently of metadata about the audio input signal ( metadata can happen in very early stage, therefore it is related to the audio input signal, see [0005]: “… conventional audio systems frequently depend on metadata tagged by content creators to guide the adjustment of audio settings”; the ML based content detection can perform without involving metadata, therefore the confidence score, i.e. the content indicator used by the ML based content detection module can be provided independently of metadata, see [0074]: “…with the disclosed techniques, automatic adaptation of audio settings for dynamically changing audio scenes are possible without metadata tagging”)
Regarding claim 5, Pramod teaches:
the content-based features are extracted in a frame-wise approach (features are content based and extracted from audio frames, therefore a frame-wise approach, see “Audio scene analysis based on audio content type identification …segmenting the audio data into a plurality of audio frames; extracting one or more audio features from the plurality of audio frames…”, Abstract), and
wherein the content indicator categorizes the audio input signal in one of the following categories: music, spoken word, audio for video, gaming audio, or miscellaneous audio (Classifier is in Audio scene analyzer, which classifies or categories, dialogue i.e. spoken word, music, environmental noise i.e. miscellaneous audio, see [0023]: “ Audio input (not shown) can… include various types of audio scenes such as dialogue, music, environmental noise, and/or the like. Audio scene analyzer 114 processes audio data 116 to determine audio scene type”; see [0042]:”FIG. 3 is a block diagram of audio type classifier 210 included in audio scene analyzer 114 processing extracted features 204…”).
Regarding claim 11, Pramod teaches at least one electro-acoustic transducer and a microphone each coupled with the at least one processor, wherein the audio input signal is configured for output by the at least one electro-acoustic transducer (Loudspeaker(s) is at least one electro-acoustic transducer, a separate media device can be a microphone, both of which are connected to the processor via the I/O interface on the Computing Device , see [0025] : “I/O interface 106 facilitates communication between computing device 102 and other external systems including the loudspeaker(s) 104. For example, audio input can be received through the I/O interface 106, such as from a separate media device (not shown)…”, See Fig. 1; processed audio input signal is outputted by the loudspeaker(s), see [0063]: ”The loudspeaker(s) 106 convert the process audio data 116 into sound waves…”).
Regarding claim 12, since the claimed method comprises the same operations conducted by the device in claim 1, claim 12 is rejected as being anticipated for the reasons mentioned in claim 1’s 102 rejection.
Regarding claim 13, since the claimed method comprises the same operations conducted by the device in claim 2, claim 13 is rejected as being anticipated for the reasons mentioned in claim 2’s 102 rejection.
Regarding claim 16, since the claimed method comprises the same operations conducted by the device in claim 5, claim 16 is rejected as being anticipated for the reasons mentioned in claim 5’s 102 rejection.
Claim Rejections - 35 USC § 103
Claims 3 and 14 are rejected under 35 U.S.C. 103 as being obvious over Pramod et al. (WO 2024145444 A1 ).
Regarding claim 3, Pramod teaches all the claim elements previously stated in claim 2’s 102 rejection.
Pramod is silent about whether or not the processor has access to the metadata about the audio input signal. However, there are two limited choices (same as the application’s specification see [0052]: “in some cases, the controller 50 does not have access to the metadata 260 about the audio input signal 110. In some additional cases, the controller 50 has access to the metadata 260 about the audio input signal 110”), i.e. the processor has access to the metadata, or the processor does not have access to the metadata. At the time of invention was effectively filed, it would have been obvious for a designer to pick one that suits their needs the most, and neither of these choices produces unexpected results.
Regarding claim 14, since the claimed method comprises the same operations conducted by the device in claim 3, claim 14 is rejected as being obvious for the reasons mentioned in claim 3’s 103 rejection.
Claims 4 and 15 are rejected under 35 U.S.C. 103 as being obvious over Pramod et al. (WO 2024145444 A1 ) in view of Querze, III et al. (US 20240029755 A1).
The applied reference has a common inventor with the instant application. Based upon the earlier effectively filed date of the reference, it constitutes prior art under 35 U.S.C. 102(a)(2).
This rejection under 35 U.S.C. 103 might be overcome by: (1) a showing under 37 CFR 1.130(a) that the subject matter disclosed in the reference was obtained directly or indirectly from the inventor or a joint inventor of this application and is thus not prior art in accordance with 35 U.S.C.102(b)(2)(A); (2) a showing under 37 CFR 1.130(b) of a prior public disclosure under 35 U.S.C. 102(b)(2)(B); or (3) a statement pursuant to 35 U.S.C. 102(b)(2)(C) establishing that, not later than the effective filing date of the claimed invention, the subject matter disclosed and the claimed invention were either owned by the same person or subject to an obligation of assignment to the same person or subject to a joint research agreement. See generally MPEP § 717.02.
Regarding claim 4, Pramod teaches all the claim elements previously stated in claim 2’s 102 rejection.
Pramod does not teach the metadata about the audio input signal is used to verify the content indicator.
Querze, III teaches the metadata about the audio input signal is used for verify the content indicator (metadata associated with the audio input is used for analysis, i.e. verification under certain conditions, which can include the content indicator exceeds a threshold; therefore, the metadata can be used to verify the content indicator, see [0113]: “the analyzing is performed using a trained machine-learning model. In aspects, metadata associated with the content is analyzed wherein the metadata includes that the content includes speech. In aspects, when one or more predefined conditions exceed a threshold, a voice track of the audio signal is analyzed”).
At the time of the invention was effectively filed, it would have been obvious to one of ordinary skill in the art to have incorporated the content indicator verification with the metadata as taught by Querze, III in device as taught by Pramod. It would have yielded predictable results and resulted in an improved device. One of ordinary skill in the art would have been motivated to do so to “enables a user to better understand the linguistic meanings” (Querze, III: [0005]).
Regarding claim 15, since the claimed method comprises the same operations conducted by the device in claim 4, claim 15 is rejected as being obvious for the reasons mentioned in claim 4’s 103 rejection.
Claims 6 and 17 are rejected under 35 U.S.C. 103 as being obvious over Pramod et al. (WO 2024145444 A1 ) in view of Nguyen et al. (“DCASE 2018 Task 2: Iterative Training, Label Smoothing, And Background Noise Normalization For Audio Event Tagging”).
Regarding claim 6, Pramod teaches all the claim elements previously stated in claim 1’s 102 rejection.
Pramod does not teach the ML based content detection module is trained using normalized training data, wherein the normalized training data is fed back into the ML based content detection module iteratively during the training.
Nguyen teaches the ML based content detection module is trained using normalized training data, wherein the normalized training data is fed back into the ML based content detection module iteratively during the training (background noise is normalized, see “we propose to use background noise normalization.”, Abstract; the normalization is part of the training, therefore, normalized noise data is at least part of the normalized training data, the iterative training process means normalized training data is fed back iteratively, see “The second model use background noise normalized STFT to extract log-mel features. We combine the three models using their geometric mean which scores higher on the PLB compared to their arithmetic mean. We start the iterative training process with the verified labels”, pg. 3, Sec. 3).
At the time of the invention was effectively filed, it would have been obvious to one of ordinary skill in the art to have normalized the training data and fed it iteratively during training as taught by Nguyen in the ML based content detection module as taught by Pramod. It would have yielded predictable results and resulted in an improved device. One of ordinary skill in the art would have been motivated to do so for “automatic label verification and label smoothing to reduce the over-fitting” (Nguyen: Abstract).
Regarding claim 17, since the claimed method comprises the same operations conducted by the device in claim 6, claim 17 is rejected as being obvious for the reasons mentioned in claim 6’s 103 rejection.
Claims 7 and 18 are rejected under 35 U.S.C. 103 as being obvious over Pramod et al. (WO 2024145444 A1 ) in view of Lu et al. (US 20180068670 A1).
Regarding claim 7, Pramod teaches all the claim elements previously stated in claim 1’s 102 rejection.
Pramod does not teach:
the processor is further configured to perform at least one of:
a) down-sampling the audio input signal prior to extracting the content-based features, or
b) applying a hysteresis factor prior to adjusting the audio output setting for the audio output to mitigate undesirable switching.
Lu teaches applying a hysteresis factor prior to adjusting the audio output setting for the audio output to mitigate undesirable switching (buffer scheme and larger thresholds compared with previous values are a hysteresis factor to prevent undesirably frequent switching, see [0460]: “…in the case where the VoIP speech or non-VoIP speech confidence is close to and fluctuates around the threshold, the VoIP/non-VoIP classification results are possible to switch too frequently. To avoid such fluctuation, a buffer scheme may be provided: both thresholds for VoIP speech and non-VoIP speech may be set larger, so that it is not so easy to switch from the present content type to the other content type”).
At the time of the invention was effectively filed, it would have been obvious to one of ordinary skill in the art to have integrated the switching scheme as taught by Lu prior to adjusting the audio output setting for the audio output as taught by Pramod. It would have yielded predictable results and resulted in an improved device. One of ordinary skill in the art would have been motivated to do so for “improving experience of audience” (Lu: Abstract).
Regarding claim 18, since the claimed method comprises the same operations conducted by the device in claim 7, claim 18 is rejected as being obvious for the reasons mentioned in claim 7’s 103 rejection.
Claims 8 and 19 are rejected under 35 U.S.C. 103 as being obvious over Pramod et al. (WO 2024145444 A1 ) in view of Sharma et al. (WO 2020209951 A1).
Regarding claim 8, Pramod teaches all the claim elements previously stated in claim 1’s 102 rejection.
Pramod does not teach the ML based content detection module is a computationally limited version of a counterpart ML based content detection module trained remotely from the audio device.
Sharma teaches the ML based content detection module is a computationally limited version of a counterpart ML based content detection module trained remotely from the audio device (A ML based model/module is the counterpart, trained remotely from the audio device, edge-ified model is the resources-constrained i.e. computationally limited ML module, see “A machine learning model is created and trained in the remote network… the model is edge-converted ("edge-ified") to run optimally with the constrained resources of the edge device and with the same or better level of accuracy. The "edge-ified" model is adapted to operate on continuous streams of sensor data in real-time and produce inferences. The inferences can be used to determine actions to take in the local network without communication to the remote network”, Abstract).
At the time of the invention was effectively filed, it would have been obvious to one of ordinary skill in the art to have adopted a computationally limited version of a ML training remotely as taught by Sharma to the detection module as taught by Pramod. It would have yielded predictable results and resulted in an improved device. One of ordinary skill in the art would have been motivated to do so to “prevent costly machine failures or downtime as well as improve the efficiency and safety of industrial operations” (Sharma: [75]).
Regarding claim 19, since the claimed method comprises the same operations conducted by the device in claim 8, claim 19 is rejected as being obvious for the reasons mentioned in claim 8’s 103 rejection.
Claim 9 is rejected under 35 U.S.C. 103 as being obvious over Pramod et al. (WO 2024145444 A1) in view of Blewett et al. (US 20070192104 A1).
Regarding claim 9, Pramod teaches all the claim elements previously stated in claim 1’s 102 rejection.
Pramod does not teach at least one processor is a fixed point device processor.
Blewett teaches at least one processor is a fixed point device processor (Processor without a FPU is a fixed point device processor, see [0026]: “a fixed-point number representation in computing is a real data type for a number that has a fixed number of digits after the decimal binary or radix) point. Fixed-point numbers are useful for representing fractional values in native two's complement format if the executing processor has no floating point unit (FPU) or if fixed-point provides improved performance. Most low cost embedded processors do not have an FPU”).
At the time of the invention was effectively filed, it would have been obvious to one of ordinary skill in the art to have used a fixed point processor as taught by Blewett to the device as taught by Pramod. It would have yielded predictable results and resulted in an improved device. One of ordinary skill in the art would have been motivated to do so for “improved performance” (Blewett: [0026]).
Claims 10 and 20 are rejected under 35 U.S.C. 103 as being obvious over Pramod et al. (WO 2024145444 A1 ) in view of Driscoll (US 20200210832 A1)
Regarding claim 10, Pramod teaches all the claim elements previously stated in claim 1’s 102 rejection.
Pramod does not teach the ML based content detection module is sized based on computational parameters of the processor.
Driscoll teaches configurations can be sized based on computational parameters of the processor (neural network model can include the ML based content detection model, identifying hardware constraints module based on computational parameters of the processor, e.g. speed, see [0041]: “…the constraints module 130 identifies one or more model constraints specified by the hardware platform for the neural network model, specific to the neural network model being implemented on the hardware platform. Example hardware constraints may relate to a processing resource of the hardware platform, such as memory size, cache size, processor information (e.g., speed), instructions capable of being implemented, and so on”; configurations can be optimized and sized according to the constraints, see [0056]: “…reduce the number of valid configurations and lead to a smaller number of potentially more optimized configurations that satisfy all constraints.”).
At the time of the invention was effectively filed, it would have been obvious to one of ordinary skill in the art to have adopted sizeable configurations based on processor’s computational parameters as taught by Driscoll to the ML based content detection module as taught by Pramod. It would have yielded predictable results and resulted in an improved device. One of ordinary skill in the art would have been motivated to do so “to implement and/or configure neural networks on previously-unimplemented platforms” (Driscoll: [0005]).
Regarding claim 20, since the claimed method comprises the same operations conducted by the device in claim 10, claim 20 is rejected as being obvious for the reasons mentioned in claim 10’s 103 rejection.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHIN LEE whose telephone number is (571)272-1460. The examiner can normally be reached Monday thru Friday 8-5 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vivian Chin can be reached at 571-272-7848. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHIN LEE/Examiner, Art Unit 2695
/VIVIAN C CHIN/Supervisory Patent Examiner, Art Unit 2695