DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claims 5, 12, and 19 are each objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 3, 8, 10, 15, and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over a combination of Panabaker et al. (US 2006/0161835) and Valmiki et al. (US 6975324).
Regarding claim 1, Panabaker teaches a method, applied to a scenario in which a source terminal projects a video frame and an audio frame to a sink terminal, wherein the method is performed by the sink terminal (Fig. 2), and the method comprises:
decoding a video frame in a video queue and an audio frame in an audio queue based on a video interval or an audio interval (Abstract, “A source of audiovisual content transmits corresponding digital data to one or more endpoints, such as over a home network, where it may be buffered and/or decoded for playback.” [0035], “The other endpoint 214, which may also be referred to as a sink device, and which may provide a remote audio and/or video display device as its output mechanism or mechanisms 220, such as a networked television set including speakers, includes a receiver 222 that receives the transmitted encoded data and places that data into a decoding buffer 224.” Fig. 2),
to obtain a decoded video frame and a decoded audio frame that meet an audio-video synchronization condition (Abstract, “A source of audiovisual content transmits corresponding digital data to one or more endpoints, such as over a home network, where it may be buffered and/or decoded for playback.” [0035], “The other endpoint 214, which may also be referred to as a sink device, and which may provide a remote audio and/or video display device as its output mechanism or mechanisms 220, such as a networked television set including speakers, includes a receiver 222 that receives the transmitted encoded data and places that data into a decoding buffer 224.” [0036], “As described above, such a system, without more, has no way to ensure that the decoders are operating on the same data at the same time. The result is that the output mechanisms of the endpoints may be out of synchronization.” Fig. 2); and
playing the decoded video frame and the decoded audio frame ([0033], “In FIG. 2, an audiovisual (A/V) source device 202 such as the computer system 110 or consumer electronic device provides data from some type of media player 204 for output to an endpoint 206, also denoted endpoint A.” [0035], “The other endpoint 214, which may also be referred to as a sink device, and which may provide a remote audio and/or video display device as its output mechanism or mechanisms 220, such as a networked television set including speakers, includes a receiver 222 that receives the transmitted encoded data and places that data into a decoding buffer 224.”),
wherein a quantity of video data that are from the source terminal and that are buffered in the video queue is maintained within the video interval (Abstract, “A source of audiovisual content transmits corresponding digital data to one or more endpoints, such as over a home network, where it may be buffered and/or decoded for playback.” [0034], “As also represented in FIG. 2, the endpoint 206 can optionally (as represented by the dashed box) receive the data via a driver 216, or receive decoded data extracted from the encoding buffer 210 and decoded via a decoder 218 (including any decode buffer).” [0035], “The other endpoint 214, which may also be referred to as a sink device, and which may provide a remote audio and/or video display device as its output mechanism or mechanisms 220, such as a networked television set including speakers, includes a receiver 222 that receives the transmitted encoded data and places that data into a decoding buffer 224.” [0055]), and
a quantity of audio data that are from the source terminal and that are buffered in the audio queue is maintained within the audio interval (Abstract, “A source of audiovisual content transmits corresponding digital data to one or more endpoints, such as over a home network, where it may be buffered and/or decoded for playback.” [0034], “As also represented in FIG. 2, the endpoint 206 can optionally (as represented by the dashed box) receive the data via a driver 216, or receive decoded data extracted from the encoding buffer 210 and decoded via a decoder 218 (including any decode buffer).” [0035], “The other endpoint 214, which may also be referred to as a sink device, and which may provide a remote audio and/or video display device as its output mechanism or mechanisms 220, such as a networked television set including speakers, includes a receiver 222 that receives the transmitted encoded data and places that data into a decoding buffer 224.” [0055]);
wherein the audio interval is determined based on the video interval and delay information, or the video interval is determined based on the audio interval and delay information ([0039], “The synchronization mechanism 232 may compare the actual output of the endpoint B 214 to the output of endpoint A in any number of ways. For example, the sensor 230 may detect both actual outputs from the endpoints' output mechanisms 206 and 220, such as the audio output, which if not in synchronization would be detected essentially as an echo, possibly having many seconds difference.” Fig. 2); and
wherein the delay information is determined based on a video delay and an audio delay, the video delay is a difference between times at which the source terminal and the sink terminal play a same video frame, and the audio delay is a difference between times at which the source terminal and the sink terminal play a same audio frame ([0039], “The synchronization mechanism 232 may compare the actual output of the endpoint B 214 to the output of endpoint A in any number of ways. For example, the sensor 230 may detect both actual outputs from the endpoints' output mechanisms 206 and 220, such as the audio output, which if not in synchronization would be detected essentially as an echo, possibly having many seconds difference.” [0055], “While the present invention has been primarily described with reference to the actual output of audio signals, it is also feasible to use video information. … For example, a sensor may detect a particular color pattern that is flashed on only a small corner of the screen for a time that is too brief to be noticed by a person. This detected color pattern may be matched to what it should be to determine whether the remote display was ahead of or behind the local display.” Fig. 2).
While Panabaker teaches video data and audio data, Panabaker does not expressly teach video frames and audio frames.
Valmiki teaches utilizing video frames and audio frames (Col. 12, lines 10-15, “The SDRAM controller 126 provides captured video frames to the external SDRAM.” Col. 95, line 61 to col. 96, line 2, “The audio interface module 2274 preferably detects and processes various audio frame errors.”).
In view of Valmiki’s teaching, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Panabaker to utilize video frames and audio frames in order to facilitate analysis and management of data.
Regarding claim 8, Panabaker teaches a computing device, comprising: at least one processor; and at least one memory, wherein the at least one memory stores a computer program, and the at least one processor is configured to execute the computer program stored in the at least one memory ([0023]-[0024]), causing the computing device to be enabled to perform the method of claim 1. The grounds of rejection under 35 USC §103 presented with respect to claim 1 are similarly applied to the remaining limitations of claim 8.
The grounds of rejection under 35 USC §103 presented with respect to claim 1 are similarly applied to claim 15.
Regarding claims 3, 10, and 17, the combination further teaches, wherein the audio interval is determined based on the video interval and the delay information, and the method further comprises:
when it is determined that the quantity of video frames in the video queue meets a first preset condition, reducing a decoding speed of the video frame; or when it is determined that the quantity of video frames in the video queue meets a second preset condition, increasing a decoding speed of the video frame (Panabaker: [0032], “As will be understood, there are numerous ways to implement the present invention, including positioning the sensor at various locations, modifying the data sent to an endpoint to help with pattern matching, instructing an endpoint to process its buffered data differently to effectively slow down or speed up, modifying the data sent to an endpoint or the amount of data buffered at the endpoint to essentially make it advance in its playback buffer or increase its size, gradually bringing endpoints into synchronization or doing so in a discrete hop, and so forth.” [0037], “If not synchronized, the synchronization mechanism 232 also determines whether to adjust an endpoint's playback clock and/or effectively speed up or slow down the output of one endpoint to get the system into synchronization. Note that one endpoint may move its clock backward/be slowed while the other is moved forward/sped up to achieve the same result.” [0049]).
Claim(s) 2, 4, 9, 11, 16, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over a combination of Panabaker, Valmiki, and Jordan (US 2008/0060045).
Regarding claims 2 and 16, the combination teaches the limitations specified above; however, the combination does not expressly teach: controlling a decoding speed of the video frame to maintain the quantity of video frames in the video queue within the video interval; or controlling a decoding speed of the audio frame to maintain the quantity of audio frames in the audio queue within the audio interval.
Jordan teaches controlling a decoding speed of a video frame to maintain a quantity of video frames in a video queue within a video interval; or controlling a decoding speed of an audio frame to maintain a quantity of audio frames in an audio queue within an audio interval ([0020], “the decoder 301 may receive packets of data from the IP input buffer 300 that may contain audio and video information that may have been compressed using a variant of a MPEG standard, for example.” [0023], “This rate mismatch may be the result of a difference between the encoder clock 202 (FIG. 2) frequency in the head end 100 (FIG. 1) and the STC 303 frequency. If this situation persists, the number of packets stored in the IP input buffer 300 may approach the capacity of the IP input buffer 300 resulting in an overflow condition. To prevent this situation, the CPU 304, via the PWM control, may increase the output frequency of the VCxO 302 by an incremental amount. This may in turn increase the decoding rate of the decoder 301, thus alleviating the mismatch condition.” [0024], “However, under this scenario the IP input buffer 300 may experience an underflow condition. That is, the IP input buffer 300 may not have enough packets to support subsequent decoding by the decoding buffer 301. To prevent this situation, the CPU 304, via the PWM control, may decrease the output frequency of the VCxO 302 by an incremental amount. This may in turn decrease the decoding rate of the decoder 301, thus alleviating the mismatch condition.”).
In view of Jordan’s teaching, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination to include controlling a decoding speed of the video frame to maintain the quantity of video frames in the video queue within the video interval; or controlling a decoding speed of the audio frame to maintain the quantity of audio frames in the audio queue within the audio interval. The modification would serve to enhance system efficiency and performance. The modification would thereby improve the user experience.
Regarding claim 9, the combination teaches the limitations specified above; however, the combination does not expressly teach: controlling a decoding speed of the video frame to maintain the quantity of video frames in the video queue to be maintained; or controlling a decoding speed of the audio frame to maintain the quantity of audio frames in the audio queue within the audio interval.
Jordan teaches controlling a decoding speed of a video frame to maintain a quantity of video frames in a video queue to be maintained; or controlling a decoding speed of an audio frame to maintain a quantity of audio frames in an audio queue within an audio interval ([0020], “the decoder 301 may receive packets of data from the IP input buffer 300 that may contain audio and video information that may have been compressed using a variant of a MPEG standard, for example.” [0023], “This rate mismatch may be the result of a difference between the encoder clock 202 (FIG. 2) frequency in the head end 100 (FIG. 1) and the STC 303 frequency. If this situation persists, the number of packets stored in the IP input buffer 300 may approach the capacity of the IP input buffer 300 resulting in an overflow condition. To prevent this situation, the CPU 304, via the PWM control, may increase the output frequency of the VCxO 302 by an incremental amount. This may in turn increase the decoding rate of the decoder 301, thus alleviating the mismatch condition.” [0024], “However, under this scenario the IP input buffer 300 may experience an underflow condition. That is, the IP input buffer 300 may not have enough packets to support subsequent decoding by the decoding buffer 301. To prevent this situation, the CPU 304, via the PWM control, may decrease the output frequency of the VCxO 302 by an incremental amount. This may in turn decrease the decoding rate of the decoder 301, thus alleviating the mismatch condition.”).
In view of Jordan’s teaching, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination to include controlling a decoding speed of the video frame to maintain the quantity of video frames in the video queue to be maintained; or controlling a decoding speed of the audio frame to maintain the quantity of audio frames in the audio queue within the audio interval. The modification would serve to enhance system efficiency and performance. The modification would thereby improve the user experience.
Regarding claims 4, 11, and 18, the combination further teaches,
wherein the audio interval is determined based on the video interval and the delay information (Panabaker: [0039], “The synchronization mechanism 232 may compare the actual output of the endpoint B 214 to the output of endpoint A in any number of ways. For example, the sensor 230 may detect both actual outputs from the endpoints' output mechanisms 206 and 220, such as the audio output, which if not in synchronization would be detected essentially as an echo, possibly having many seconds difference.” [0055], “While the present invention has been primarily described with reference to the actual output of audio signals, it is also feasible to use video information. … For example, a sensor may detect a particular color pattern that is flashed on only a small corner of the screen for a time that is too brief to be noticed by a person. This detected color pattern may be matched to what it should be to determine whether the remote display was ahead of or behind the local display.”).
However, the combination does not expressly teach when it is determined that the quantity of audio frames in the audio queue meets a third preset condition, reducing a decoding speed of the audio frame, or temporarily stopping decoding the audio frame; or when it is determined that the quantity of audio frames in the audio queue meets a fourth preset condition, increasing a decoding speed of the audio frame, or discarding one or more audio frames in the audio queue.
Jordan teaches: when it is determined that the quantity of audio frames in the audio queue meets a third preset condition, reducing a decoding speed of the audio frame, or temporarily stopping decoding the audio frame; or when it is determined that the quantity of audio frames in the audio queue meets a fourth preset condition, increasing a decoding speed of the audio frame, or discarding one or more audio frames in the audio queue ([0020], “the decoder 301 may receive packets of data from the IP input buffer 300 that may contain audio and video information that may have been compressed using a variant of a MPEG standard, for example.” [0023], “This rate mismatch may be the result of a difference between the encoder clock 202 (FIG. 2) frequency in the head end 100 (FIG. 1) and the STC 303 frequency. If this situation persists, the number of packets stored in the IP input buffer 300 may approach the capacity of the IP input buffer 300 resulting in an overflow condition. To prevent this situation, the CPU 304, via the PWM control, may increase the output frequency of the VCxO 302 by an incremental amount. This may in turn increase the decoding rate of the decoder 301, thus alleviating the mismatch condition.” [0024], “However, under this scenario the IP input buffer 300 may experience an underflow condition. That is, the IP input buffer 300 may not have enough packets to support subsequent decoding by the decoding buffer 301. To prevent this situation, the CPU 304, via the PWM control, may decrease the output frequency of the VCxO 302 by an incremental amount. This may in turn decrease the decoding rate of the decoder 301, thus alleviating the mismatch condition.”).
In view of Jordan’s teaching, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination to include when it is determined that the quantity of audio frames in the audio queue meets a third preset condition, reducing a decoding speed of the audio frame, or temporarily stopping decoding the audio frame; or when it is determined that the quantity of audio frames in the audio queue meets a fourth preset condition, increasing a decoding speed of the audio frame, or discarding one or more audio frames in the audio queue. The modification would serve to enhance system efficiency and performance. The modification would thereby improve the user experience.
Claim(s) 6, 13, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over a combination of Panabaker, Valmiki, and Ogawa (US 2010/0166382).
Regarding claims 6, 13, and 20, the combination teaches the limitations specified above; however, the combination does not expressly teach further teaches, further comprising: before decoding the video frame in the video queue and the audio frame in an audio queue based on the video interval or the video interval, collecting the video delay and the audio delay; and using a difference between the video delay and the audio delay as the delay information.
Ogawa teaches, before distributing content, collecting video delay and audio delay, and using a difference between the video delay and the audio delay as delay information ([0073] The synchronization control unit 28 of the server 10 calculates a difference delay time t1d of the audio delay time t1ad and the video delay time t1vd. The synchronization control unit 28 of the server 10 controls the delay time in the audio coded data storage unit 22 and the audio output data storage unit 24, or the video coded data storage unit 25 and the video output data storage unit 27 so that the difference delay time t1d becomes short. The server 10 thereby adjusts the distribution timing of the video data and the audio data to the display 11 and the audio system 12.” [0074]-[0075]).
In view of Ogawa’s teaching, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination to include before decoding the video frame in the video queue and the audio frame in an audio queue based on the video interval or the video interval, collecting the video delay and the audio delay; and using a difference between the video delay and the audio delay as the delay information. The modification would serve to enhance system efficiency and performance. The modification would thereby improve the user experience.
Claim(s) 7 and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over a combination of Panabaker, Valmiki, Jordan, and Cao et al. (US 2024/0259126).
Regarding claims 7 and 14, the combination further teaches, wherein:
when the audio interval is determined based on the video interval and the delay information,
the audio-video synchronization condition comprises
that a difference between start playing moments of an audio frame and a video frame that are synchronously played at the source terminal and that are separately played at the sink terminal is less than a first threshold, and
the first threshold is determined based on duration of an interval between two adjacent video frames; or
when the video interval is determined based on the audio interval and the delay information ([0039], “The synchronization mechanism 232 may compare the actual output of the endpoint B 214 to the output of endpoint A in any number of ways. For example, the sensor 230 may detect both actual outputs from the endpoints' output mechanisms 206 and 220, such as the audio output, which if not in synchronization would be detected essentially as an echo, possibly having many seconds difference.” Fig. 2).
However, the combination does not expressly teach the audio-video synchronization condition comprises that a difference between start playing moments of an audio frame and a video frame that are synchronously played at the source terminal and that are separately played at the sink terminal is less than a second threshold, and the second threshold is determined based on duration of an interval between two adjacent audio frames.
Jordan teaches an audio-video synchronization condition comprises that a difference between start playing moments of an audio frame and a video frame that are synchronously played at a source terminal and that are separately played at a sink terminal is less than a threshold ([0038], “The optimal delay may be found by comparing the amount of data in the IP input buffer 300 (FIG. 3) to the various threshold levels 400, 401, 402, and 406 (FIG. 4) and adjusting the STC 303 in response to results from the comparing. The time it takes to adjust the STC may be decreased by using a lower overflow threshold during, for example, a calibration mode and by using a higher overflow threshold during, for example, a post-calibration mode of operation. Oscillations in the rate of change of the STC 303 may be prevented by slew limiting the rate of change.”).
In view of Jordan’s teaching, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination wherein the audio-video synchronization condition comprises that a difference between start playing moments of an audio frame and a video frame that are synchronously played at the source terminal and that are separately played at the sink terminal is less than a second threshold. The modification would serve to facilitate determination of conditions requiring synchronization processing.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify that the second threshold is determined based on duration of an interval between two adjacent audio frames.
Cao provides a teaching for duration of an interval between two adjacent audio frames ([0054], “The video module obtains the synchronization anchor time information of the instance corresponding to the kernel driver through the method of C++ library, decides output time and strategy of the current video frame, so as to realize the synchronous output of the video module and the audio module. For example, system reference clocks corresponding respectively to audio timestamps of two adjacent frames are determined so as to obtain a system reference clock deviation, time when a next audio frame is played (that is, time when the upcoming audio is played) is determined based on the synchronization anchor time and the system reference clock deviation, and the video frame output is controlled at the time point corresponding to the time when the upcoming audio is played, so as to achieve audio and video synchronization output.”).
In view of Cao’s teaching, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination such that the second threshold is determined based on duration of an interval between two adjacent audio frames. The modification would serve to aid in accurate synchronization of content.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Hui et al. (US 2020/0177947) discloses a method that improves synchronization between a video source and a video sink during a video streaming process ([0041], Fig. 4).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL R TELAN whose telephone number is (571)270-5940. The examiner can normally be reached 9:30AM-6:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Nasser Goodarzi can be reached at (571) 272-4195. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL R TELAN/ Primary Examiner, Art Unit 2426