Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on May 26, 2026 has been entered.
Status of Claims
Claims 1, 2, 6, 7, 9, 12, 13, 15 and 16 are amended
Claims 1 – 16 remain pending.
Response to Arguments
Applicant's arguments filed 05/26/2026 with respect to claims 1 – 16 have been considered but are moot because the new grounds of rejection do not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 7, 10, 12 – 16 are rejected under 35 U.S.C 103 as being unpatentable over Chakraborty et al. US Patent Publication No. US-20160070963-A1 (hereinafter Chakraborty) in view of Townsend US Patent Application Publication No. US-9620168-B1 (hereinafter Townsend) and further in view of Fujino et. al. US Patent Application Publication No. US-20160116737-A1 (hereinafter Fujino) and Song Patent Application Publication No. CN-105678702-A (hereinafter Song).
Regarding claim 1, Chakraborty discloses a method performed by a computing device comprising: performing a frame score generation process to calculate a frame diversity score (Chakraborty in [0037] discloses, “In the exemplary embodiment illustrated in FIG. 2B, frame-scoring logic 235 includes both coverage scoring logic 230 and diversity scoring logic 240”), calculating a plurality of scores (In [0037] Chakraborty discloses about calculating the diversity score of the frame), calculating a plurality of aesthetic diversity scores, wherein each aesthetic diversity score of the plurality is based on a pairwise aesthetic feature difference between scene-related features of a respective frame of the set of frames and scene-related features of the first frame (In [0037] Chakraborty discloses about determining diversity score and in [0036] it discloses about object detection (aesthetic) across different frames); and a respective aesthetic diversity score (Chakraborty in [0037] discloses about determining diversity score and in [0036] it discloses about object detection (aesthetic) across different frames);
Chakraborty doesn’t disclose the limitations as further recited in the claim.
Townsend discloses receiving a stream of image data defining a first frame and a set of frames not including the first frame (Townsend in [Column – 10, 35 – 37] discloses, “As illustrated in FIG. 4, the server(s) 112 may receive (410) video data and may optionally receive (412) existing annotation data associated with the video data”. In Fig. 5A Townsend discloses about determining about first frame and not including first frame); and determining, using a minimal frame diversity score of the plurality of frame diversity scores, whether to include the first frame as part of an image object representing suggested frames of the stream of image data (Townsend in [Column – 3, Line 29 – 42] discloses, “To select from the candidate video clips, the server(s) 112 may determine similarity scores between the candidate video clips. For example, the server(s) 112 may determine a first similarity score between a first candidate video clip and a second candidate video clip .... For example, the characteristics of the candidate video clips may include feature vectors, histograms of image data, gradients of the image data, histograms of gradients, a signature of the image data or the like that may be used to determine if image data is similar between candidate video clips”).
It would been obvious to one with one having an ordinary skill in art before the effective filling date of the claimed invention to integrate the technique of Townsend into the system of Chakraborty because it would allow the system to suggest frames that not only distinct by facial expression but also captured in different time frame.
Chakraborty and Townsend in the combination doesn’t disclose the limitations as further recited in the claim.
Fujino discloses the frame score generation process comprising: calculating a plurality of time diversity scores, wherein each time diversity score of the plurality is based on a pairwise time feature difference between time-related features of a respective frame of the set of frames and time-related features of the first frame (Fujino in [0059] discloses, “The time-axis direction similarity determination unit 906 determines the similarity between frames having different times by using the feature points and feature amounts received by the second data reception unit. The position and orientation measurement unit 907 measures a position and orientation using the similarity, feature points, and feature amounts, and outputs position and orientation information”), and calculating a plurality of frame diversity scores, wherein each frame diversity score of the plurality is based on a respective time diversity score (Fujino in [0059] discloses about feature difference on different time frame, “The time-axis direction similarity determination unit 906 determines the similarity between frames having different times by using the feature points and feature amounts received by the second data reception unit”).
It would have been obvious to one of ordinary skill in art before the effective filling date of the claimed invention to integrate the technique of Fujino into the system of Chakraborty and Townsend because it would allow the system to avoid suggesting redundant frame that were captured to too close in time and select frames that are more diverse.
Chakraborty, Townsend and Fujino in the combination doesn’t disclose the limitations as further recited in the claim.
Song discloses each facial diversity score of the plurality is based on a pairwise facial feature difference between facial-related features of a respective frame of the set of frames and facial-related features of the first frame, ); a respective facial diversity score (Song in [Page – 3, Paragraph – 1] discloses, “the sending end module comprises a face feature point detecting unit and feature point sending unit, a face feature point detecting unit for collecting each frame camera according to the time sequence of the human face image sequence frame coordinate information of the face feature point detection. and the calculating difference of face feature points detected by the feature point data of the coordinate information and the initial image, the above feature point variation value output by the preset change threshold T, feature point sending unit for characteristic point change value of each frame image according to the time sequence output by the face feature point detecting unit”).
It would have been obvious to one of ordinary skill in art before the effective filling date of the claimed invention to integrate the technique of Song into the system of Chakraborty, Townsend and Fujino because it would allow the system to select frames that contain diverse or improved facial expression.
Summary of Citations (Townsend)
[Column – 3, Line 29 – 42]; “To select from the candidate video clips, the server(s) 112 may determine similarity scores between the candidate video clips. For example, the server(s) 112 may determine a first similarity score between a first candidate video clip and a second candidate video clip .... For example, the characteristics of the candidate video clips may include feature vectors, histograms of image data, gradients of the image data, histograms of gradients, a signature of the image data or the like that may be used to determine if image data is similar between candidate video clips”.
[Column – 3, Line 52 – 65 & Column – 4, Line 1 – 4]; “the server(s) 112 may select any of the candidate video clips having a peak priority metric value ... In some examples, the server(s) 112 may select the first video clips based on priority metrics using a variable threshold .... Thus, the server(s) 112 may group the candidate video clips based on the similarity scores and may select a desired number of candidate video clips from each group based on a highest priority metric. For example, the server(s) 112 may group ten candidate video clips together based on the similarity score and may select three candidate video clips having the highest priority metric as the first video clips to increase a diversity between the first video clips”.
[Column – 9, Line 32 – 34]; “the server(s) 112 may analyze a video frame 310 and generate annotation data 312 , which may include time (e.g., a timestamp, a period of time, etc.)”.
[Column – 10, Line 35 – 37]; “As illustrated in FIG. 4, the server(s) 112 may receive (410) video data and may optionally receive (412) existing annotation data associated with the video data”.
[Column – 12, Line 43 – 46]; “As illustrated in FIG. 5B, the annotation database 512 includes Frame 1, Frame 2, Frame 3, Frame 10, Frame 11, Frame 30 and Summary Data associated with the overall video clip”.
[Column – 17, Line 9 – 12]; “In some examples, the server(s) 112 may associate different priority metrics with an object over time, such as when a face is obscured, hidden and/or turned away from the image capture device 110”.
Summary of Citations (Chakraborty)
Paragraph [0036]; “the feature vector may include features determined using any object detection technique known in the art. In the exemplary embodiment, frame feature extractor 229 is to generate a feature vector comprising histograms of oriented gradient (HOG) features”.
Paragraph [0037]; “In the exemplary embodiment illustrated in FIG. 2B, frame-scoring logic 235 includes both coverage scoring logic 230 and diversity scoring logic 240”.
Summary of Citations (Fujino)
Paragraph [0059]; “The time-axis direction similarity determination unit 906 determines the similarity between frames having different times by using the feature points and feature amounts received by the second data reception unit. The position and orientation measurement unit 907 measures a position and orientation using the similarity, feature points, and feature amounts, and outputs position and orientation information”.
Summary of Citations (Song)
[Page – 3, Paragraph – 1]; “the sending end module comprises a face feature point detecting unit and feature point sending unit, a face feature point detecting unit for collecting each frame camera according to the time sequence of the human face image sequence frame coordinate information of the face feature point detection. and the calculating difference of face feature points detected by the feature point data of the coordinate information and the initial image, the above feature point variation value output by the preset change threshold T, feature point sending unit for characteristic point change value of each frame image according to the time sequence output by the face feature point detecting unit”.
Regarding claim 7, Song in the combination discloses the method of claim 6, wherein calculating the plurality of facial diversity scores for the frames of the set of frames relative to the first frame further comprises (Song in [Page – 3, Paragraph – 1] discloses, “the sending end module comprises a face feature point detecting unit and feature point sending unit, a face feature point detecting unit for collecting each frame camera according to the time sequence of the human face image sequence frame coordinate information of the face feature point detection. and the calculating difference of face feature points detected by the feature point data of the coordinate information and the initial image”).
Townsend further discloses about extracting facial features from the selected frame and from the first frame, the extracted facial features representing at least one of facial features depicted in the frames or face embeddings (Townsend in [Column – 10; Line 53 – 64] discloses about identifying facial features, “the server(s) 112 may analyze the video frame and identify the face(s) based on facial recognition, identifying head and shoulders, identifying eyes, smile recognition or the like”. And in Fig. 5A it discloses about identifying first frame); and determining a facial feature difference between the selected frame and the first frame utilizing the extracted facial features and the distance metric (Townsend in [Column – 14; Line 1 – 3] discloses, “As illustrated in FIG. 7, the server(s) 112 may determine a similarity score between individual video frames of a portion of video data” wherein calculating the similarity implies to calculating the distance metrics).
Summary of Citations (Townsend)
[Column – 10; Line 60– 63]; “the server(s) 112 may analyze the video frame and identify the face(s) based on facial recognition, identifying head and shoulders, identifying eyes, smile recognition or the like. Optionally, the server(s) 112 may determine (420) identities associated with the face(s)”.
[Column – 14; Line 1 – 3]; “As illustrated in FIG. 7, the server(s) 112 may determine a similarity score between individual video frames of a portion of video data”.
Summary of Citations (Song)
[Page – 3, Paragraph – 1]; “the sending end module comprises a face feature point detecting unit and feature point sending unit, a face feature point detecting unit for collecting each frame camera according to the time sequence of the human face image sequence frame coordinate information of the face feature point detection. and the calculating difference of face feature points detected by the feature point data of the coordinate information and the initial image”.
Regarding claims 10, the combination of Chakraborty, Townsend, Fujino and Song as a whole teaches claim 1, and Chakraborty teaches claim 10 for the same grounds of rejection from the Final Office Action of 01/28/2026.
Regarding claim 12, Townsend in the combination discloses the method of claim 1, wherein calculating each frame diversity score of the plurality for the frames of the set of frames relative to the first frame based on the facial diversity score, the aesthetic diversity score, and the time diversity score comprises: computing a weighted sum of the respective facial diversity score, the respective aesthetic diversity score (Townsend in Fig. 5A discloses about comparing time, Faces and object (aesthetic) between frames).
Fujino discloses the respective time diversity score (Fujino in [0059] discloses, “The time-axis direction similarity determination unit 906 determines the similarity between frames having different times by using the feature points and feature amounts received by the second data reception unit”).
Summary of Citations (Fujino)
Paragraph [0059]; “The time-axis direction similarity determination unit 906 determines the similarity between frames having different times by using the feature points and feature amounts received by the second data reception unit”.
Regarding claim 13, Townsend in the combination discloses the method of claim 1, wherein calculating the plurality of time diversity scores for frames of the set of frames relative to the first frame comprises: determining a timestamp difference between a selected frame of the set of frames and the first frame; and generating a time diversity score based on the determined timestamp difference (Townsend in Fig. 5A discloses about comparing time between frames).
Regarding claims 14, the combination of Chakraborty, Townsend, Fujino and Song as a whole teaches claim 1, and Townsend teaches claim 14 for the same grounds of rejection from the Final Office Action of 01/28/2026.
Regarding claim 15, apparatus claim 15 corresponds to method claim 1. Therefore, the rejection analysis and motivation to combine of claim 1 is applicable to claim 15.
Regarding claim 16, is a computer readable storage medium claim corresponds to method claim 1. Therefore, the rejection analysis of claim 1 is applied in claim 16. See also in [0087].
Summary of Citations (Chakraborty)
Paragraph [0087]; “In one or more third embodiment, a computer-readable storage media has instructions stored...”.
Claims 2 – 6, 8, 9 and 11 are rejected under 35 U.S.C 103 as being unpatentable over Chakraborty in view of Townsend, Fujino and Song and further in view of Jiang US Patent Application Publication No. US-10134440-B2 (hereinafter Jiang).
Regarding claim 2, Chakraborty in the combination discloses the method of claim 1, wherein determining, using the minimal frame diversity score of the plurality of frame diversity scores, whether to include the first frame as part of an image object representing suggested frames of the stream of image data further comprises: storing, in a ring buffer, the set of frames (Chakraborty in [0040] discloses, “Processed video data may be buffered in a FIFO manner, for example with a circular buffer”); iteratively performing, for each stored frame (Chakraborty disclosed in [0045]), a filtering process comprising: selecting a frame from the stored frames (Chakraborty discloses in [0045]; “Candidate frames are selected based on the solution that optimizes the frame scoring for the selection. At operation 495 , the batch of k stream summary frames is updated in response to the selection of frames made at operation 440 differing from the batch of k stream summary frames accessed at operation 406”); assigning a drop score to the selected frame; determining the stored frame with a lowest drop score; evicting the stored frame with the lowest drop score from the ring buffer; and storing, in the ring buffer, the first frame (Chakraborty in [0054] discloses, “for each frame stored at operation 450 , the corresponding coverage score is stored in association with the frame ... One non-selected incumbent frame is removed/replaced in the stream summary for each non-incumbent frame selected. Any coverage score associated with the non-selected incumbent frame is also removed/replaced” wherein the frame with the lowest score is removed).
Chakraborty and Townsend, Fujino and Song in the combination doesn’t disclose the limitations as further recited in the claim.
Jiang discloses calculating a frame quality score for the selected frame (Jiang in [Column – 10, Line 38 – 39] discloses, “model for determine overall image quality scores 440”).
It would been obvious to one with one having an ordinary skill in art before the effective filling date of the claimed invention to integrate the technique of Jiang into the system of Chakraborty in view of Townsend, Fujino and Song because it would help in selecting frames that are both diverse with high quality frames.
Summary of Citations (Chakraborty)
Paragraph [0040]; “Processed video data may be buffered in a FIFO manner, for example with a circular buffer”.
Paragraph [0045]; “In response to receiving each new frame set (e.g., V, V+1, etc.) a summarization iteration is performed where the incumbent k stream summary frames”.
Paragraph [0046]; “Each non-incumbent frame in a new set of frames received from the CM is scored with respect to the other frames in the new set, and with respect to each incumbent frame. Candidate frames are selected based on the solution that optimizes the frame scoring for the selection. At operation 495 , the batch of k stream summary frames is updated in response to the selection of frames made at operation 440 differing from the batch of k stream summary frames accessed at operation 406”.
Paragraph [0053]; “The top k entries in the solution vector may then be set to 1 and the remainder set to 0, to reconstruct the integer solution. This optimization vector identifies the non-incumbent frames and incumbent frames to be discarded (0 valued elements) and selected as summary frames (1 valued elements)”.
Paragraph [0054]; “for each frame stored at operation 450 , the corresponding coverage score is stored in association with the frame to facilitate a subsequent comparison with a next set of frames (e.g., to be read in at 506 in FIG. 5B). One non-selected incumbent frame is removed/replaced in the stream summary for each non-incumbent frame selected. Any coverage score associated with the non-selected incumbent frame is also removed/replaced”.
Summary of Citations (Jiang)
[Column – 10, Line 38 – 39]; “model for determine overall image quality scores 440”.
Regarding claims 3 – 5, Chakraborty in view of Townsend, Fujino and Song fails to teach the limitations as recited in claims 3 – 5 respectively. However, Jiang does. The grounds of rejection and motivation to combine from Final Office Action 01/28/2026 with respect to Jiang apply here.
Regarding claim 6, Song in the combination discloses the method of claim 1, wherein calculating the plurality of facial diversity scores for frames of the set of frames relative to the first frame comprises: determining facial-related features for the first frame (Song in [Page – 3, Paragraph – 1] discloses, “the sending end module comprises a face feature point detecting unit and feature point sending unit, a face feature point detecting unit for collecting each frame camera according to the time sequence of the human face image sequence frame coordinate information of the face feature point detection. and the calculating difference of face feature points detected by the feature point data of the coordinate information and the initial image”).
Jiang further discloses about iteratively performing a scoring process comprising: selecting a frame from the set of frames; determining facial-related features for the selected frame; and utilizing a distance metric to determine a facial feature difference between the facial-related features of the selected frame and the facial-related features of the first frame (Jiang in [Column – 12; Line 1 – 7] discloses about the distance metric between frames, “A select key image frames step 470 is used to select the set of key image frames 280 from the candidate key image frames 465. For each image frame subset 405 visual “distances” between the candidate key image frames 465 are determined based on differences between the corresponding visual feature vectors”. And [Column – 11; Line 23 -26] discloses that the system determine facial related features).
Summary of Citations (Song)
[Page – 3, Paragraph – 1]; “the sending end module comprises a face feature point detecting unit and feature point sending unit, a face feature point detecting unit for collecting each frame camera according to the time sequence of the human face image sequence frame coordinate information of the face feature point detection. and the calculating difference of face feature points detected by the feature point data of the coordinate information and the initial image”.
Summary of Citations (Jiang)
[Column – 11; Line 23 – 26]; “In addition to measuring the overall image quality, the assess visual quality step 410 also determines facial quality scores 420 for each image frame 210 in the image frame subsets 405”.
[Column – 12; Line 1 – 7]; “A select key image frames step 470 is used to select the set of key image frames 280 from the candidate key image frames 465. For each image frame subset 405 visual “distances” between the candidate key image frames 465 are determined based on differences between the corresponding visual feature vectors”.
Chakraborty in view of Townsend, Fujino and Song fails to teach the limitations as recited in claim 8 respectively. However, Jiang does. The grounds of rejection and motivation to combine from Final Office Action 01/28/2026 with respect to Jiang apply here.
Regarding claim 9, Chakraborty in the combination discloses the method of claim 1, wherein calculating the plurality of aesthetic diversity scores for frames of the set of frames relative to the first frame comprises: determining scene-related features for the first frame; and iteratively performing a scoring process comprising: selecting a frame from the set of frames; extracting scene-related features from the selected frame; and utilizing a distance metric to determine an aesthetic feature difference between the scene-related features of the selected frame and the scene-related features of the first frame (Jiang in [Column – 11; Line 10 – 14] discloses, “the spatial distribution of high-frequency edges, the color distribution, the hue entropy, the blur degree, the color contrast, and the brightness (6 dimensions) are computed as visual features” wherein determining the blur degree equates to determining aesthetic. And in [Column – 12; Line 1 – 7] Jiang discloses about calculating distances of multiple frames equates to determining plurality of diversity score).
Regarding claim 11, Chakraborty in view of Townsend, Fujino and Song fails to teach the limitations as recited in claim 11 respectively. However, Jiang does. The grounds of rejection and motivation to combine from Final Office Action 01/28/2026 with respect to Jiang apply here.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner
should be directed to ZAID MUHAMMAD SALEH whose telephone number is (703)756-1684.
The examiner can normally be reached M-F 8 am - 5 pm ET. Examiner interviews are available
via telephone, in-person, and video conferencing using a USPTO supplied web-based
collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO
Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.If attempts to
reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be
reached on (571)272-7332. The fax phone number for the organization where this application or
proceeding is assigned is 571-273-8300. Information regarding the status of published or
unpublished applications may be obtained from Patent Center. Unpublished application
information in Patent Center is available to registered users. To file and manage patent
submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit
https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center
and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For
additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
If you would like assistance from a USPTO Customer Service Representative, call 800-786
9199 (IN USA OR CANADA) or 571-272-1000.
/ZAID MUHAMMAD SALEH/
Examiner, Art Unit 2668
06/27/2026
/VU LE/Supervisory Patent Examiner, Art Unit 2668