DETAILED ACTION
This communication is in response to the Application filed on 12/10/2024. Claims 1-7, 17-18 and 20-30 are pending and have been examined. Claims 1, 17 and 18 are independent. Claims 8-16 and 19 are canceled.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 03/05/2025 and 02/17/ 2026 were filed. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Priority
This application is a 371 of PCT/CN2023/129685 submitted on 11/03/2023.
Acknowledgment is made of applicant’s claim for foreign priority based on application CN202211387602.8 filed in China National Intellectual Property Administration (CNIPA) on 11/07/2022 and receipt of a certified copy thereof.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 17-18, 20 and 26 are rejected under 35 U.S.C. 103 as being unpatentable over Ishikawa et al. (US Pub 2013/0090929) in view of Zhan et al. (US Pat 8,510,121).
Regarding Claim 1,
Ishikawa discloses an audio data encoding method (Fig.20, par [065], "… a block diagram illustrating a framework of a low delay hybrid encoder having two encoding modes..."), comprising:
determining a coding mode of a first audio frame (Fig.20, par [073], "…The incoming signal is also sent to a signal classification block 2003. The signal classification decides which coding mode is selected for a time domain signal...");
judging whether the coding mode of the first audio frame is a same coding mode as a coding mode of a second audio frame (Fig.20, par [0], "…A mode indicator from the signal classification block 2003 is sent to the bit multiplexer block 2006...");
wherein the second audio frame is a previous audio frame of the first audio frame; in response to the coding mode of the first audio frame being different from the coding mode of the second audio frame (Fig.2, par [005], "…the hybrid codec needs to have block switching methods for transition frames in which the coding mode switches...the current frame is defined as a transition frame."; par [075], "…In FIG. 20, the block switching algorithms 2002 are used to handle the transition frames where the coding mode is switched..." ) and
the coding mode of the first audio frame is multiple description coding, generating third data based on first data, second data and a first delay (Fig.4, par [076], "…The block switching algorithm concatenates the second half of the previous frame i-1 to form an extended frame...");
the first data is low-frequency data obtained by frequency division of original audio data of the first audio frame, the second data is low-frequency data obtained by frequency division of original audio data of the second audio frame, and the first delay is a coding delay of the multiple description coding (Fig.20, par [073], "…The current time domain signal in low frequency band to be coded is sent to a corresponding encoder2004, 2005 according to the mode indicator….");
Ishikawa does not explicitly discloses the limitations, "the coding mode of the first audio frame is multiple description coding," and "performing the multiple description coding on the third data to obtain encoded data of the first audio frame."
Zhan, in the analogous field of a multiple description audio coding in communication, discloses performing the multiple description coding on the third data to obtain encoded data of the first audio frame (Zhan, Fig2b, Fig.3, col.4:16-65, "…human ears are sensitive to a low-frequency part and less sensitive to a high-frequency part. Therefore, considering speech quality and bit rate redundancy, a low-frequency part obtained by dividing the residual signals may be coded by using a multiple description method with good speech quality...").
Therefore, it would have been obvious to a person of ordinary skill in the art to apply the known frequency-band-division and MDC technique of Zhan to the code-switching boundary architecture of a hybrid encoder/decoder system of Ishikawa to improve the device with a reasonable expectation that this would result in a hybrid encoder/decoder device that could perform coding or decoding while switching between SDC and MDC codecs and also enhance the quality of audio transmission using the frequency division technique to reduce the bit rate of MDC coding/decoding. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings Ishikawa and Zhan to obtain the invention as specified in claim 1.
Regarding Claim 2,
Ishikawa in view of Zhan discloses the method of claim 1, wherein the method further comprises: in response to that the coding mode of the first audio frame is different from that of the second audio frame (Ishikawa, Fig.20, par [073], "…The incoming signal is also sent to a signal classification block 2003. The signal classification decides which coding mode is selected for a time domain signal..."; par [075], "…In FIG. 20, the block switching algorithms 2002 are used to handle the transition frames where the coding mode is switched...") and
the coding mode of the first audio frame is single description coding (Ishikawa, par [072], "…the ACELP only uses a sample of the current frame, i.e., one frame, to code the current frame…"; i.e., ACLEP is the single frame mode ion steady state, which is structurally analogous to SDC),
generating sixth data based on fourth data, fifth data and a second delay (Ishikawa, Fig.4, par [076], "…The block switching algorithm concatenates the second half of the previous frame i-1 to form an extended frame..."; i.e., "the added length" from extending the frame is a delay);
the fourth data is the original audio data of the first audio frame, the fifth data is the original audio data of the second audio frame, and the second delay is a coding delay of the single description coding (Ishikawa, the coding mode of the first audio frame is a target mode. ACLEP is a target mode and "the added length" from extending the frame is a delay); and
performing the single description coding on the sixth data to obtain encoded data of the first audio frame (Fig.20, par [073], "…The current time domain signal in low frequency band to be coded is sent to a corresponding encoder 2004, 2005 according to the mode indicator….").
Claim 17 is a device claim with limitations similar to the limitations of Claim 1 and is rejected under similar rationale. Additionally,
Zhan discloses an electronic device comprising: a memory and a processor, wherein the memory is used for storing a computer program; the processor is used to, when executing the computer program, cause the electronic device to (Zhan, col.10:4-39, Embodiment 5).
…
Rationale for combination is similar to that provided for Claim 1.
Claim 18 is a non-transitory computer-readable storage medium claim with limitations similar to the limitations of Claim 1 and is rejected under similar rationale. Additionally,
Zhan discloses a non-transitory computer-readable storage medium having a computer program stored thereon, the computer program, when executed by a computing device, causes the computing device to implement (Zhan, col.10:4-39, Embodiment 5).
…
Rationale for combination is similar to that provided for Claim 1.
Claim 20 is a device claim with limitations similar to the limitations of Claim 2 and is rejected under similar rationale.
Claim 26 is a non-transitory computer-readable storage medium claim with limitations similar to the limitations of Claim 2 and is rejected under similar rationale.
Claim 3-4, 21-22, and 27-28 are rejected under 35 U.S.C. 103 as being unpatentable over Ishikawa in view of Zhan further in view of Smithers et al. (US Pub 2005/0102049).
Regarding Claim 3,
Ishikawa in view of Zhan discloses the method of claim 1, wherein the generating the third data based on the first data, the second data and the first delay, which is an ad hoc "concatenate then encode" approach and lacks an windowed smoothing disclosed in Claim 3.
Smithers, in the analogous field of audio signal processing, especially splicing frame-based audio information, discloses intercepting samples with length of the first delay from a tail end of the second data to obtain seventh data (Smithers, Fig.7a, par [110], "…the faded-down appendage an' is shown appended to the end of the (n-1)st frame to provide a frame with faded appendage 70...");
splicing the seventh data at a head end of the first data to obtain eighth data (Fig.7b, par [113], "…at the System output, in a concatenation of formerly non-sequential frames n-1 and n+1 at a frame boundary 76 (i.e., audio frames that have a discontinuity). The crossfaded portion provides a smooth crossfade from the PCM audio in frame n-1 to that in frame n+1..."); and
deleting samples with the length of the first delay from the tail end of the eighth data to obtain the third data (Fig.7b, par [113], "…An appendage such as appendage an+2, appended to the end of frame n+1 may be removed, if it is not associated with another crossfade...").
Therefore, it would have been obvious to a person of ordinary skill in the art to apply Smithers' crossfade technique in a hybrid mode-switching codec device taught by Ishikawa in view of Zhan to improve the device with a reasonable expectation that this would result in an codec device eliminating audible discontinuity at a frame splice point for the predictable result of a smooth SDC/MDC transition (Smithers, paras ([005-011]).
Claim 4 is a method claim with limitations similar to the limitations of Claim 3 and is rejected under similar rationale. Rationale for combination is similar to that provided for Claim 3.
Claim 21 is a device claim with limitations similar to the limitations of Claim 4 and is rejected under similar rationale.
Claim 22 is a device claim with limitations similar to the limitations of Claim 3 and is rejected under similar rationale.
Claim 27 is a non-transitory computer-readable storage medium claim with limitations similar to the limitations of Claim 4 and is rejected under similar rationale.
Claim 28 is a non-transitory computer-readable storage medium claim with limitations similar to the limitations of Claim 3 and is rejected under similar rationale.
Claims 5, 23 and 29 are rejected under 35 U.S.C. 103 as being unpatentable over Ishikawa in view of Zhan further in view of Li et al. (US Pub 2016/0165060).
Regarding Claim 5,
Ishikawa in view of Zhan discloses the method of claim 1, but does not explicitly teaches the determining the coding mode of the first audio frame.
Li, in the analogous field of endeavor, discloses determining whether a coding mode switching condition is met based on a signal type of the first audio frame and a coding mode duration (Li, Fig.27, paras [163-169], "…At 2706, the method 2700 includes, monitoring, during the communication session, an operational condition of the communication network…At 2708, the method 2700 includes, deciding to switch, when the operational condition of the communication network meets a first condition...");
wherein the coding mode duration is a playback duration of an audio frame continuously encoded in a current coding mode (Li, par [102], "…whereby no change is made when a previous codec change occurred within a preceding threshold interval of time (i.e., playback duration)...") ;
In response to the coding mode switching condition being not met, determining the coding mode of the second audio frame as the coding mode of the first audio frame (Li, Fig.13, par [102], "…The method 1300 may further include operating the user device to maintain a history of codec changes, and applying a hysteresis in the determination about whether or not to change the current audio codec or the current video codec..."; see also codec switching rule book in Fig.26); and
in response to the coding mode switching condition being met, determining the coding mode of the first audio frame according to network parameters of an encoded audio data transmission network (Li, Fig.27, par [168], "…At 2710, the method 2700 includes, switching, after deciding to switch, transmission of the first media content of the communication session to use the second media codec technology...").
Therefore, it would have been obvious to a person of ordinary skill in the art to apply Li's known hysteresis/threshold switching duration time technique in a hybrid mode-switching codec device taught by Ishikawa in view of Zhan to improve the device with a reasonable expectation that this would solve the stability problem in adaptive codec switching for the predictable result of a stabilized switching decision to satisfy user expectation with low latency and high fidelity interaction (Li, paras ([001-003]).
Claim 23 is a device claim with limitations similar to the limitations of Claim 5 and is rejected under similar rationale.
Claim 29 is a non-transitory computer-readable storage medium claim with limitations similar to the limitations of Claim 5 and is rejected under similar rationale.
Claims 6, 24 and 30 are rejected under 35 U.S.C. 103 as being unpatentable over Ishikawa in view of Zhan further in view of Li further in view of Sehlstedt (US Pub 2012/0215536).
Regarding Claim 6,
Ishikawa in view of Zhan further in view of Li discloses the method of claim 5, wherein the determining whether the coding mode switching condition is met based on the signal type of the first audio frame and the coding mode duration, comprises: judging whether the coding mode duration is greater than a threshold duration (Li, par [102], "…whereby no change is made when a previous codec change occurred within a preceding threshold interval of time (i.e., playback duration)..."), but does not explicitly teaches the coding mode switching decision based on probability of audio frame being voice or non-voice.
Sehlstedt, in the analogous field of threshold adaptation for the voice activity detector, discloses judging whether a probability that the first audio frame is a voice audio frame is less than a threshold probability (Sehlstedt, Fig.1, par [003], "…an overview block diagram of a generalized VAD 180...a VAD decision 160 is a decision for each frame whether the frame contains speech or noise"; par [004], "…A primary decision, "vad_prim" 150, is made by a primary voice activity detector 140 and is basically just a comparison of the features for the current frame and the background features estimated from previous input frames, where a difference larger than a threshold causes an active primary decision...An operation controller 110 may adjust the threshold(s) for the primary detector and the length of the hangover according to the characteristics of the input signal..."; Fig.2, paras [043-045], "…the processor 203 is configured to detect whether the received frame comprises voice based on the comparison between the first SNR and the adaptive threshold..." );
in response to the coding mode duration being greater than the threshold duration and the probability that the first audio frame is the voice audio frame being less than the threshold probability, determining that the coding mode switching condition is met; and in response to the coding mode duration being less than or equal to the threshold duration and/or the probability that the first audio frame is the voice audio frame being greater than or equal to the threshold probability, determining that the coding mode switching condition is not met.
Therefore, it would have been obvious to a person of ordinary skill in the art to apply Sehlstedt's known threshold compared voice-activated-detection technique, a well-established method in the same field for distinguish voice from non-voice content, in a hysteresis-gated codec device taught by Ishikawa in view of Zhan further in view of Li to produce the predictable result of the combined duration and voice probability switching gate to provide a VAD with improved performance (Sehlstedt, paras ([006-014]).
Claim 24 is a device claim with limitations similar to the limitations of Claim 6 and is rejected under similar rationale.
Claim 30 is a non-transitory computer-readable storage medium claim with limitations similar to the limitations of Claim 6 and is rejected under similar rationale.
Claims 7 and 25 are rejected under 35 U.S.C. 103 as being unpatentable over Ishikawa in view of Zhan further in view of Li further in view of Sehlstedt further in view of Leung (US Pat 11,916,664).
Regarding Claim 7,
Ishikawa in view of Zhan further in view of Li discloses the method of claim 5, but does not explicitly teaches the determining the coding mode based on the packet loss rate of the encoded audio data transmission network according to the network parameter.
Leung, in the analogous field of adjusting a configuration of an audio coder-decoder (codec), discloses wherein the determining the coding mode of the first audio frame according to the network parameters of the encoded audio data transmission network, comprises: determining a packet loss rate of the encoded audio data transmission network according to the network parameters (Leung, Fig.1, col.5:15-57, "…system 100 includes a configuration server 142 communicatively coupled via a network 150 to one or more devices...The codec configuration adaptation circuitry 137 is configured to, in response to receipt of the request 192, update the codec configuration 116 based on the configuration data 107...The codec configuration adaptation circuitry 137 is configured to determine a packet loss rate (PLR) 129..."; Fig.5, col.13:5-49, "…The method 500 also includes determining a packet loss rate at the first device, at 504...");
judging whether the packet loss rate is greater than or equal to a threshold packet loss rate (Leung, Fig.1, col.5:15-57, "…based on determining that a decoder of the first device has the first codec configuration and that the packet loss rate satisfies the first packet loss rate threshold, sending, to the second device, a request to change a codec configuration of the second device, at 506...");
in response to the packet loss rate being greater than or equal to the threshold packet loss rate, determining that the coding mode of the first audio frame is the multiple description coding; and in response to the packet loss rate is less than the threshold packet loss rate, determining that the coding mode of the first audio frame is single description coding.
Ishikawa in view of Zhan further in view of Li further in view of Sehlstedt discloses the switching condition gate (duration + voice probability) and the crossfade transition mechanism, but no step for determining the target configuration from the network condition.
Therefore, it would have been obvious to a person of ordinary skill in the art to apply Leung's packet loss rate and threshold technique, which selects more robust codec configuration (functionally analogous to MDC since the redundancy is the art recognized for loss robustness) above the threshold and the less robust (i.e. SDC), in a voice-gated adaptive codec device taught by Ishikawa in view of Zhan further in view of Li further in view of Sehlstedt to complete a network adaptive configuration determination procedure as a predictable result.
Claim 25 is a device claim with limitations similar to the limitations of Claim 7 and is rejected under similar rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Beack et al. (US Pat 8,954,321) discloses a Unified Speech and Audio Codec (USAC) that may process a window sequence based on mode switching is provided. The USAC may perform encoding or decoding by overlapping between frames based on a folding point when mode switching occurs (Beack, section).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JANGWOEN LEE whose telephone number is (703)756-5597. The examiner can normally be reached Monday-Friday 8:00 am - 5:00 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, BHAVESH MEHTA can be reached at (571)272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JANGWOEN LEE/ Examiner, Art Unit 2656
/BHAVESH M MEHTA/ Supervisory Patent Examiner, Art Unit 2656