DETAILED ACTION
This communication is in response to the Application filed on 12/27/2024. Claims 1-20 are pending and have been examined. Claims 1 and 11 are independent. This Application was published as US20260188329A1.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The abstract of the disclosure is objected to because Abstract exceeds 150 words (153 words). A corrected abstract of the disclosure is required and must be presented on a separate sheet, apart from any other text. See MPEP § 608.01(b).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-4, 7-8, 10-11, 13-14, 17-18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Valin et al. (US Pat 11,924,367) in view of Vaillancourt et al. (US Pub 2018/0233154) further in view of Tahernezhaadi et al. (US Pat 6,785,339).
Regarding Claim 1,
Valin discloses a computerized method for downlink audio processing (Valin, Fig.1, audio enhancement system 100) comprising:
receiving an audio signal at an uplink device; converting the audio signal to a near-end signal (Fig.1, col.5:1-23, "…near-end communication device 110 and far end communication device 120. Each device may respectively implement sensors for capturing audio data, such as microphones 112 and 122, and loudspeakers or other playback devices for providing transmitted audio data, such as speakers 114 and 124..."; Fig.6, col.12:14-36, "…at 610, first audio data may be received (e.g., near audio data) captured by a microphone at a first communication device...");
collecting a far-end signal (col.12:37-45, "…at 620, second audio data (e.g., far-end audio data) transmitted from the second communication device to the first communication device as part of the two-way communication...");
Valin discloses "performing audio processing" (e.g. an acoustic echo canceller 104, echo and noise suppression 106, and deep neural network (DNN) model 427 for audio enhancement (Fig.4)) in a two-way communication 101, and transmitting the enhanced audio signal via network, but does not explicitly disclose the rest of limitations in Claim 1.
Vaillancourt, in the analogous field of stereo sound encoding, discloses encoding the near-end signal and the far-end signal into an encoded bit stream (Vaillancourt, Fig.1, par [045], "…A stereo sound encoder 106 encodes the left 105 and right 125 channels of the digital stereo sound signal thereby producing a set of encoding parameters that are multiplexed under the form of a bitstream 107...");
transmitting the encoded bit stream to a downlink device (Fig.1, par [046], encoded bitstream 111 is transmitted to a stereo decoder across a communication link 101.);
decoding the encoded bit stream at the downlink device to render a decoded near-end signal and a decoded far-end signal (par [046], "…A stereo sound decoder 110 converts the received encoding parameters in the bitstream 112 for creating synthesized left 113 and right 133 channels of the digital stereo sound signal…");
Vaillancourt discloses an audio-coding that jointly encodes two related channel signals into a single multiplexed bitstream using a stereo codec, while Valin's two-way audio communication system treats near-end and far-end signals as a related pair. Therefore, It would have been obvious to a person of ordinary skill in the art to apply Vaillancourt's joint stereo-bitstream encoding technique in Valin's near-end/far-end signal pairs to predictably yield a single combined bitstream that carries both signals using less transmission bandwidth with a reasonable expectation that the quality and the intelligibility of stereo speech is greatly improved for complex audio scene (Vaillancourt, paras [005, 039-0040]).
Valin in view of Vaillancourt does not explicitly teach the limitation, "performing audio processing at the downlink device using the decoded near-end signal and the decoded far-end signal to generate processed audio."
Tahernezhaadi, in the analogous field of speech enhancement in communication systems, discloses performing audio processing at the downlink device using the decoded near-end signal and the decoded far-end signal to generate processed audio (Tahernezhaadi, Fig.5, col.9:16-col.10:64, "…the network element (e.g., gateway) 504 performs speech decoding on received packets from both directions. The decoded speech samples are then fed into the adaptive speech quality detector 514 to check in parallel for the presence of echo, noise or poor audio level...", "…The packet quality enhancer 515 provide echo cancellation, volume level control, noise suppression, vocoder bypass as known in the art, in response...").
Therefore, It would have been obvious to a person of ordinary skill in the art to apply Tahernezhaadi's "decode-then-enhance" technique to the audio enhancement system with joint-stereo encoding/decoding taught by Valin in view of Vaillancourt to predictably yield the system with a reasonable expectation that the further combined system would produce the enhanced near-end and far-end signals right after two channels are recovered from Vaillancourt's stereo decoder "at receiving end unit" (Tahernezhaadi, Background).
Regarding Claim 3,
Valin in view of Vaillancourt further in view of Tahernezhaadi discloses the method of claim 1, wherein encoding the near-end signal and the far-end signal utilizes an audio codec with stereo coding capabilities (Vaillancourt, Fig.1, par [045], "…A stereo sound encoder 106 encodes the left 105 and right 125 channels of the digital stereo sound signal thereby producing a set of encoding parameters that are multiplexed under the form of a bitstream 107..."; (par [046], "…A stereo sound decoder 110 converts the received encoding parameters in the bitstream 112 for creating synthesized left 113 and right 133 channels of the digital stereo sound signal...").
Regarding Claim 4,
Valin in view of Vaillancourt further in view of Tahernezhaadi discloses the method of claim 3, wherein the codec includes at least one of Opus stereo, AAC, EVS, and IVAS (Vaillancourt, par [040], "…The latest 3GPP EVS conversational speech standard provides a bit-rate range from 7.2 kb/s to 96 kb/s for wideband (WB) operation and 9.6 kb/s to 96 kb/s for super wideband (SWB) operation...").
Regarding Claim 7,
Valin in view of Vaillancourt further in view of Tahernezhaadi discloses the method of claim 1, wherein the audio processing includes acoustic echo cancellation (Valin, col.4:58-65, "…audio enhancement system 100 may implement an acoustic echo canceller 104...and echo noise suppression 106, similar to the discussion above and below with regard to FIG.4..."; Tahernezhaadi, Fig.5, col.9:16-col.10:64, "…The packet quality enhancer 515 provide echo cancellation, volume level control, noise suppression, vocoder bypass as known in the art, in response...").
Regarding Claim 8,
Valin in view of Vaillancourt further in view of Tahernezhaadi discloses the method of claim 7, wherein the audio processing further includes adaptive noise suppression and automatic gain control (Tahernezhaadi, Fig.5, col.9:16-col.10:64, "…The packet quality enhancer 515 provide echo cancellation, volume level control, noise suppression, vocoder bypass as known in the art, in response...").
Regarding Claim 10,
Valin in view of Vaillancourt further in view of Tahernezhaadi discloses the method of claim 1, further comprising performing additional processing of the processed audio using at least one machine learning algorithm (Valin, col.9:24-col.10:19, "…Deep neural network model 427 may generate ideal ratio masks, which are provided as to envelope post-filter 429...").
Claim 11 is a system claim with limitations similar to the limitations of Claim 1 and is rejected under similar rationale. Rationale for combination is similar to that provided for Claim 1.
Claim 13 is a system claim with limitations similar to the limitations of Claim 3 and is rejected under similar rationale.
Claim 14 is a system claim with limitations similar to the limitations of Claim 4 and is rejected under similar rationale.
Claim 17 is a system claim with limitations similar to the limitations of Claim 7 and is rejected under similar rationale.
Claim 18 is a system claim with limitations similar to the limitations of Claim 8 and is rejected under similar rationale.
Claim 20 is a system claim with limitations similar to the limitations of Claim 10 and is rejected under similar rationale.
Claims 2 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Valin in view of Vaillancourt further in view of Tahernezhaadi further in view of Ojala et al. (US Pub 2008/0114606).
Regarding Claim 2,
Valin in view of Vaillancourt further in view of Tahernezhaadi discloses the method of claim 1, but does not explicitly discloses a jitter butter at the downlink device.
Ojala, in the analogous field of processing multi-channel audio signals, discloses wherein the downlink device includes a jitter buffer when decoding the encoded bit stream (Ojala, Figs.3-5, paras [032-040], "…The methods for adaptive jitter management may be used for adding or removing short segments of signal within an audio frame of downmixed multi-channel audio signal. The basic idea is to re-synchronize the side information to the time scaled downmixed audio...The method is especially suitable for speech communication applications with spatial audio...").
Therefore, it would have been obvious to one of ordinary skill in the art, before effective filing date of the claimed invention, to have modified the audio enhancement system for jointly encoded dual channel signal of Valin in view of Vaillancourt further in view of Tahernezhaadi with Ojala's adaptive jitter buffering technique and time-scaling method with a reasonable expectation of success to smooth out packet-arrival irregularities before decoding and audio enhancement pipeline by dynamically controlling the balance between short enough delay and low enough number of delayed frames (Ojala, paras [002-007]).
Claim 12 is a system claim with limitations similar to the limitations of Claim 2 and is rejected under similar rationale. Rationale for combination is similar to that provided for Claim 2.
Claims 5-6 and 15-16 are rejected under 35 U.S.C. 103 as being unpatentable over Valin in view of Vaillancourt further in view of Tahernezhaadi further in view of Krueger et al. ("A New Approach for Low-Delay Joint-Stereo Coding," ITG Conference on Voice Communication [8. ITG-Fachtagung], Aachen, Germany, 2008, pp. 1-4.).
Regarding Claim 5,
Valin in view of Vaillancourt further in view of Tahernezhaadi discloses the method of claim 1, but does not explicitly discloses the limitation, "wherein encoding the near-end signal and the far-end signal utilizes a mono coder that interleaves the near-end signal and the far-end signal in a time domain."
Krueger, in the analogous field of a stereo coding, discloses wherein encoding the near-end signal and the far-end signal utilizes a mono coder that interleaves the near-end signal and the far-end signal in a time domain (Krueger, Abstract, "…a new approach for the coding of stereophonic audio signals based on inter-channel linear prediction is proposed..."; Fig.2, 3 The New Approach, "…Our new approach operates in the time domain to achieve low algorithmic delay and is shown in Figure 2. From the right and the left channel input signal, in the first step a mono signal is calculated...the mono signal xM(k), the two sets of (N + 1) stereo prediction coefficients adi) and aR(i) and the residual signals eL(k) and eR(k) are quantized and transmitted...").
Therefore, it would have been obvious to one of ordinary skill in the art, before effective filing date of the claimed invention, to have modified the audio enhancement system for jointly encoded dual channel signal of Valin in view of Vaillancourt further in view of Tahernezhaadi with a time-domain joint-stereo coding technique based on linear prediction techniques of Krueger with a reasonable expectation of success to extend any existing monaural speech or audio codec such as G.711 or G.722 toward stereo functionality and achieve low algorithmic delay in time-domain, higher SNR, and improved perceived stereo audio quality (Krueger, Conclusion).
Regarding Claim 6,
Valin in view of Vaillancourt further in view of Tahernezhaadi further in view of Krueger disclose method of claim 5, wherein the mono coder includes G.711 or G.722 (Krueger, Abstract, "…Due to its modularity, it is also suitable to extend any existing monaural speech or audio codec toward stereo functionality..."; "suitable to extend any existing monaural audio codec…" directly reads on G.711/G.722).
Claim 15 is a system claim with limitations similar to the limitations of Claim 5 and is rejected under similar rationale. Rationale for combination is similar to that provided for Claim 5.
Claim 16 is a system claim with limitations similar to the limitations of Claim 6 and is rejected under similar rationale.
Claims 9 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Valin in view of Vaillancourt further in view of Tahernezhaadi further in view of Seguin (US Pub 2009/0253457).
Regarding Claim 9,
Valin in view of Vaillancourt further in view of Tahernezhaadi discloses the method of claim 1, but does not explicitly discloses "playing back the processed audio at the downlink device."
Seguin, in the analogous field of an audio signal processing in wireless communication, discloses further comprising playing back the processed audio at the downlink device (Fig.3: a receiver (ear speaker or earpiece) 220, and a speaker (speakerphone) 218, par [034], "…speakers 218, 220, which generates a speaker driver signal whose strength is based on an adjustable volume setting received as input..."; par [035], "…The downlink audio processor 08 may perform several digital signal processing operations upon the decoded, digital or sampled voice signal or bitstream (also referred to as the downlink baseband or audio signal)...operations performed on the downlink audio signal: adjust its gain using a downlink programmable gain amplifier (PGA) 310; apply general filtering to it...perform multi-band audio compression...reduce noise using a noise suppressor 316…").
Therefore, it would have been obvious to a person of ordinary skill in the art to apply the Seguin's handheld communication device architecture houses codec, audio enhancement, and local playback together in the combined audio enhancement system of Valin in view of Vaillancourt further in view of Tahernezhaadi with a reasonable expectation that the same downlink device outputting processed audio through its own speaker.
Claim 19 is a system claim with limitations similar to the limitations of Claim 9 and is rejected under similar rationale. Rationale for combination is similar to that provided for Claim 9.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Boland (US Pat 7,936,705) discloses multiplexing methods, devices, and systems that provide the ability to playback multiple VoIP audio streams simultaneously with a single RTP session and further provides the ability to conference all streams together prior to transmission (Boland, Abstract).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JANGWOEN LEE whose telephone number is (703)756-5597. The examiner can normally be reached Monday-Friday 8:00 am - 5:00 pm ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, BHAVESH MEHTA can be reached at (571)272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JANGWOEN LEE/ Examiner, Art Unit 2656
/BHAVESH M MEHTA/ Supervisory Patent Examiner, Art Unit 2656