Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Election/Restrictions
Applicant’s election without traverse of Claims 1 and 13-20 in the reply filed on 5/18/2026 is acknowledged. Claims 2-12 are withdrawn from further consideration pursuant to 37 CFR 1.142(b) as being drawn to a nonelected invention, there being no allowable generic or linking claim at this time.
Allowable Subject Matter
Claims 17-18 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 13, 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Kordasiewicz (US Patent 9,191,284) in view of Winkler (US 2011/0169963).
Regarding Claim 1, Kordasiewicz (US Patent 9,191,284) discloses a method (detecting stream-switch events, Column 19 lines 40-45) comprising:
obtaining, by the processing system (media client 104 receives media content, a signal STREAMING_MEDIA through network 110, Column 7 lines 40-55), a first recorded frame (first frame of STREAMING_MEDIA, inferred from Column 7 lines 40-55) of a first variant (e.g., a first operating point, inferred from Column 7 lines 15-30) of a plurality of variants (all available operating points, Column 7 lines 15-30) associated with the reference copy of the video (absolute highest quality level, Column 27 line 1), wherein the plurality of variants comprises a plurality of copies of the video encoded at different bitrates (adaptive bit rate streaming based on video operating points characteristics such as bit rate, Column 5 lines 35-45).
Kordasiewicz does not disclose, but Winkler (US 2011/0169963) teaches, a method (video quality measurement system of Figs. 1A -1B implementing the three steps of temporal alignment of videos, spatial scaling, and detecting luminance gains and offsets by minimizing MSE between blocks from the two videos [0018]-[0019]) comprising:
obtaining (reference video file is stored in storage unit 120 [0020], [0022], Fig. 1A, 1B), by a processing system including at least one processor (computer system 100, a conventional computer having a processing unit [0019]), a scaled version (pyramid search [0027]) of a reference copy of a video (a reference video file 115 stored in storage unit 120 [0020]), comprising a plurality of scaled versions (pyramid search [0027]) of a plurality of frames (+/- 30 frames [0027]) of the reference copy of the video (frames [0020]);
obtaining, by the processing system (video decoder 112, which is an instantiation of computer 100, decodes the test video stream from the video data and decoded frames from the test video stream is stored in test video frame buffer 114 [0022]], Fig. 1B), a first recorded frame of a first variant (decoded frames from the test video stream is stored in test video frame buffer 114 [0022]], Fig. 1B) … associated with the reference copy of the video (reference video is used to measure the quality of the test video, Abstract);
generating, by the processing system, a first scaled version (pyramid search [0027]) of the first recorded frame (recently received frame of video A [0038]);
obtaining, by the processing system, a reference frame index (frame N in video A [0025]; initial temporal shift [0027]; frame X+N in video B with temporal shift of X [0025], reference video stream B [0038]) for a processed recorded frame (recently received [0038]) of the first variant (video A [0025], the most recently received frame of video A [0038]) … associated with the reference copy of the video (detect various changes in the spatial and temporal structure, as well as luminance and color, between two videos, referred to herein as video A and video B, both videos have the same content in terms of scene [0017]);
selecting, based on the reference frame index (initial temporal shift [0027], and frame No. N [0025]), a search set (search space [0027]) comprising a moving window (maximum horizontal and vertical shift of +/- 16 pixels [0027]; search for a block in the search space [0026]-[0027]; best match of two blocks in the search space [0026]) of the plurality of scaled versions (pyramid search [0027]) of the plurality of frames (search range for temporal shift +/- 30 frames [0027]) of the reference copy of the video (reference video stream B [0038]), the moving window being centered (+/- 30 frames [0027], +/- 16 pixels [0027]) on, or positioned relative to (initial spatial shift, initial temporal shift [0027]), a frame of the reference copy of the video (reference video stream B [0038]) having the reference frame index (frame N [0025] adjusted by initial temporal shift [0027]);
calculating, by the processing system (alignment module 130 of computer 100 performs… [0021], Figs. 1A-1B), a first plurality of image distances (pixel differences, e.g., mean square error [0026], [0018]) between the first scaled (pyramid search [0027]) version of the first recorded frame (video A, [0026]) and respective scaled versions (pyramid search [0027]) of the frames of the reference copy of the video (+/- 30 frames [0027] of the frame No. N [0025] and the initial temporal shift in video B [0027] for blocks at the initial spatial shift and within +/- 16 pixels in vertical and horizontal distance [0027]) included in the search set (temporal search space, e.g., +/- 30 frames [0027]);
and determining, by the processing system, a first frame index (Frame No. N [0025]) of the first recorded frame (test video stream [0022]) in accordance with a first least image distance (matches are based on pixel differences, e.g., mean square error; the best match is the one with the lowest mean square error [0026]) from among the first plurality of image distances that is calculated (across the search space [0026]; search space constraints in [0027]).
One of ordinary skill in the art before the application was filed would have been motivated to use the video alignment method of Winkler to measure quality of experience in Kordasiewicz because Winkler teaches that video alignment is needed for video quality measurement, and Winkler’s method can be performed continuously in real time (Abstract), improving efficiency of Kordasiewicz’s analysis system.
Regarding Claim 13, Kordasiewicz (US Patent 9,191,284) discloses the method of claim 1.
Kordasiewicz does not disclose, but Winkler (US 2011/0169963) teaches further comprising obtaining the reference copy of the video (computer 100 has a reference video file 115 stored in storage unit 120 [0019]-[0020]).
One of ordinary skill in the art before the application was filed would have been motivated to use the method of Winkler to analyze the quality of the stream of Kordasiewicz because Winkler teaches that video alignment is needed for video quality measurement, and Winkler’s method can be performed continuously in real time (Abstract), improving efficiency of the video testing system.
Regarding Claim 19, Kordasiewicz (US Patent 9,191,284) discloses a non-transitory computer-readable medium storing instructions which, when executed by a processing system including at least one processor, cause the processing system to perform operations (software system, Column 12 lines 30-40). The remainder of Claim 19 is rejected on the grounds provided in Claim 1.
Regarding Claim 20, Kordasiewicz (US Patent 9,191,284) discloses a device comprising:
a processing system including at least one processor (microprocessor, co-processors, a micro-controller, digital signal processor, microcomputer, central processing unit, field programmable gate array, programmable logic device, Column 18 lines 10-15);
and a computer-readable medium storing instructions which, when executed by the processing system, cause the processing system to perform operations (software system, Column 12 lines 30-40). The remainder of Claim 19 is rejected on the grounds provided in Claim 1.
Claim(s) 14-16 are rejected under 35 U.S.C. 103 as being unpatentable over Kordasiewicz (US Patent 9,191,284) in view of Winkler (US 2011/0169963) and Li (US 2015/0023404).
Regarding Claim 14, Kordasiewicz (US Patent 9,191,284) discloses the method of claim 13 further comprising: obtaining a plurality of recorded frames for at least a portion of the plurality of variants (stream switch event indicating a switch between operating points has occurred during streaming, a single, identical media session with a large variety and number of operating point changes, Column 25 line 64 – Column 26 line 20) associated with the reference copy of the video (absolute highest quality level, Column 27 line 1).
Kordasiewicz does not disclose, but Li (US 2015/0023404) teaches and calculating image distances (PSNR/MSE [0021]) between the plurality of recorded frames of the at least the portion of the plurality of variants (each quality level [0021]) and respective corresponding frames of the reference copy of the video having same frame indexes (inherent: PSNR and MSE are full-reference objective measures, therefore, a reference frame is inherently included in the disclosed method).
One of ordinary skill in the art before the application was filed would have been motivated to detect the stream_switch event in Kordasiewicz using explicit indication of quality such as PSNR, as in Li, because Li teaches that content delivery can be optimized using PSNR rather than average bitrate to inform the client of quality level [0016], improving the constant quality experience. Note that although Li does not calculate PSNR/MSE in the process of inferring quality (rather, it measures and explicitly indicates the quality), Li none-the-less renders obvious the act of inferring quality because its method is identical to the method of inferring quality. Li suggests that the PSNR /MSE of the quality levels are sufficiently distinct that they uniquely identify the quality level of the stream (i.e., “explicit indication corresponding to a quality metric such as …PSNR… MSE….”). Because there is a forward mapping from quality level to PSNR/MSE, there is also a backward mapping from PSNR/MSE to quality level.
Regarding Claim 15, Kordasiewicz (US Patent 9,191,284) discloses the method of claim 14.
Kordasiewicz does not disclose, but Li (US 2015/0023404) teaches further comprising:
calculating a first image distance between the first recorded frame (PSNR/MSE [0021]) and a first frame of the reference copy having the first frame index (inherent: PSNR and MSE are full-reference objective measures, therefore, a reference frame is inherently included in the disclosed method).
Li does not teach, but renders obvious and determining a variant of the first recorded frame in accordance with a closest match between the first image distance and one of the image distances for recorded frames of the at least the portion of the plurality of variants having the first frame index (Li renders obvious the act of inferring quality because Li’s method is identical to the method of inferring quality. Li suggests that the PSNR/MSE of the quality levels are sufficiently distinct that they uniquely identify the quality level of the stream (i.e., “explicit indication corresponding to a quality metric such as …PSNR… MSE….”). Because there is a forward mapping from quality level to PSNR/MSE, there is also a backward mapping from PSNR/MSE to quality level. This renders inferring quality form PSNR obvious to those of ordinary skill in the art.).
One of ordinary skill in the art before the application was filed would have been motivated to detect the stream_switch event in Kordasiewicz using explicit indication of quality such as PSNR, as in Li, because Li teaches that content delivery can be optimized using PSNR rather than average bitrate to inform the client of quality level [0016], improving the constant quality experience. Note that although Li does not calculate PSNR/MSE in the process of inferring quality (rather, it measures and explicitly indicates the quality), Li none-the-less renders obvious the act of inferring quality because its method is identical to the method of inferring quality. Li suggests that the PSNR /MSE of the quality levels are sufficiently distinct that they uniquely identify the quality level of the stream (i.e., “explicit indication corresponding to a quality metric such as …PSNR… MSE….”). Because there is a forward mapping from quality level to PSNR/MSE, there is also a backward mapping from PSNR/MSE to quality level.
Regarding Claim 16, which repeats the steps of Claim 15 on a second frame, is rejected on the grounds provided in Claim 15.
Response to Arguments
Applicant’s remarks filed 4 September 2026 have been considered but are not persuasive.
Applicant’s remarks regarding Winkler, Remarks at 10-11, are not persuasive. Applicant argues that Winkler does not teach the claimed moving window of scaled reference-frame versions selected based on a reference frame index for a processed recorded frame from an adaptive bitrate variant. Remarks at 10. Applicant continues, the disclosure does not describe or suggest the claimed operation of obtaining a reference frame index for a processed recorded frame of an ABR variant and using that index to select a moving window comprising scaled versions of reference-video frames. Remarks at 10. The cited teaching does not disclose that the search set is selected based on a reference frame index for a processed recorded frame, that the search set comprises a moving window of scaled versions of reference-video frames, or that image distances are calculated between a scaled recorded frame and the respective scaled reference frames included in that search set. Remarks at 10-11.
Examiner agrees that Winkler does not apply its real-time, full-reference video alignment for quality testing to adaptive bitrate video, but Applicant’s argument regarding the application of Winkler to adaptive bitrate video is not persuasive because the rejection is based on a combination of Kordasiewicz and Winkler, in which the system of Winkler is incorporated into the adaptive bitrate testing of Kordasiewicz. Thus, Applicant’s argument regarding the application of Winkler to ABR video is unpersuasive as it argues the Winkler alone when the rejection is based on a combination.
Winkler does teach a “search space comprising a moving window of the plurality of scaled versions of the plurality of frames of the reference copy of video.” Specifically, Winkler teaches that an example block size for block matching is 64x64 pixels [0027], and the search space is +16 pixels and -16 pixels in both horizontal and vertical directions around the initial spatial shift. [0027]. Within this search space, a block is “searched for” [0026], and the “best match” block is found. [0026]. Although the phrase “moving window” is not used, this description is of a moving window search inside the search space. It is standard and sufficient to connote a moving window search within the search space.
Winkler also teaches that the search space is around the reference frame index. Winkler teaches that the initial temporal shift is estimated [0027], and the final shift could be X frames before or after Frame No. N of the incoming frame [0025]. The search space is +/- 30 frames from the initial temporal shift. [0027]. Thus, the frame index N and the initial temporal shift are the reference frame index and the search space is +/- 30 frames around that temporally central frame index.
Winkler also teaches image distances calculated between the recorded frame and the reference frame. Winkler teaches mean-squared error used to measure the difference between the received block and blocks in the search space. [0026].
And although Winkler does not expressly state scaling the video for block matching, it does use the phrase “optimized block matching… such as pyramid search.” [0027]. Pyramid search is a standard technique involving hierarchical block matching. A resolution pyramid is formed by down-scaling the images several times, and matching is conducted from low resolution to high resolution. It is a standard technique in the art and Winkler’s terminology invokes all features inherent to it.
Bilibrov is not necessary in the rejection as its teachings are redundant with Winkler’s pyramid search. Although Hamming distance is not the same as mean squared error, it is still used to measure the difference between images, so it is an “image distance.” Applicant has not narrowed the claimed “image distance” to exclude the metric that Bilibrov uses to measure the image distance. Regardless, this argument is moot as Bilibrov is not necessary in the rejection.
Caspi (US Patent 7,428,345) describes a pyramid search and can be relied upon as evidence.
For these reasons, Applicant’s arguments are unpersuasive.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
EP 2227031 A2 – detecting skipped frames
US 20080253689 A1 – registration between test and reference video sequences
US Patent 7,428,345 – describes a pyramid search.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHADAN E HAGHANI whose telephone number is (571)270-5631. The examiner can normally be reached M-F 9AM - 5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jay Patel can be reached at 571-272-2988. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHADAN E HAGHANI/Examiner, Art Unit 2485