DETAILED ACTION
Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
2. This office action is responsive to Applicant’s Arguments/Remarks Made in an Amendment received on 06/18/2026.
Status of Claims
3. Claims 1-32 are pending in this application.
Claims 1-12 are currently presented for examination.
Claim 1 is currently amended.
Claims 13-32 were previously withdrawn from consideration.
Response to Arguments
4. Regarding Applicant’s Argument (pages 11-13):
Applicant’s arguments with respect to newly amended claim limitation(s) to independent claim 1 have been considered but are moot because the arguments are addressed by newly cited Cao reference explained in the body of rejection below. In particular newly added limitation(s) "and obtaining alignment of the detected content element through one or more transformations including translation, scale, or rotation" is met by newly cited Cao (US 2014/0185924 A1) in combination with previously cited Simhadri in view of Begun.
Applicant’s arguments concerning Simhadri paragraph [0088] are acknowledged but do not address the rejection as presently formulated. The rejection does not rely on paragraph [0088] alone for the newly added alignment limitation. Simhadri paragraphs [0038] and [0085] teach searching for keypoints throughout a bounding box, and paragraph [0038] teaches an arrangement of keypoint vertices and connecting lines constituting a geometric structure. Cao teaches landmark-based face alignment and, in paragraph [0025], computing a similarity transform to normalize a facial-landmark shape to a mean shape to obtain scale and rotation invariance. Accordingly, the combined teachings disclose or suggest the amended limitation.
Applicant’s discussion of Lakhani is also acknowledged. Lakhani is not relied upon in the rejection and remains cited as pertinent prior art of record. Therefore, Applicant’s characterization of Lakhani does not overcome the rejection over Begun, Simhadri, and Cao.
Claim Rejections - 35 USC § 103
5. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
6. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
7. The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
8. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
9. Claims 1-6, and 8-9, and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Begun et al. (US 2021/0044864) in view of Simhadri et al. (US 2021/0209734), and further in view of Cao et al. (US 2014/0185924 A1), hereinafter ‘Begun’, ‘Simhadri’ and ‘Cao’.
Regarding Claim 1:
Begun discloses a method for content recognition, the method comprising:
Begun describes a ‘method and apparatus for identifying video content based on biometric features of characters’; (Begun: ¶ [0002]).
sampling a source content for performing content recognition;
Begun discloses receiving video content and selecting or using one or more frames of the video content for analysis. More particularly, Begun receives video content at operation S510 and detects biometric features of characters from one or more frames of the video content at operation S520. The selection and use of one or more frames from the received video content constitutes sampling source content for content recognition. (Begun: Fig. 5; ¶ [0061]; see also Fig. 2 and ¶ [0058]).
detecting content elements from the sampled source content; and
Begun discloses detecting faces of actors or actresses from a frame of video content at operation S210. Begun further discloses detecting biometric features, including face shapes, from one or more frames at operation S520. The detected persons, faces, and corresponding biometric-feature-bearing regions constitute detected content elements. (Begun: Fig. 2, operation S210; Fig. 5, operation S520; ¶¶ [0058] and [0061]).
identifying the detected content elements,
Begun discloses identifying the names of actors or actresses based on their detected faces at operation S220. Begun also discloses identifying corresponding characters at operation S530 by matching detected biometric features against biometric features stored in a person-specific biometric feature database. Begun therefore identifies the detected content elements. (Begun: Fig. 2, operation S220; Fig. 5, operation S530; ¶¶ [0058; 0062-0065]).
Begun does not expressly disclose wherein detecting content elements from the sampled source content comprises: detecting content elements using an element detection model; and generating bounding boxes over each detected content element.
Simhadri discloses wherein detecting content elements from the sampled source content comprises: detecting content elements using an element detection model; and generating bounding boxes over each detected content element.
Simhadri discloses a detection component that detects a person in a frame using a trained object-detection algorithm, such as a trained YOLOv3 neural-network algorithm. The trained object-detection algorithm constitutes an element detection model. (Simhadri: Fig. 1; ¶¶ [0036-0037] and [0077]).
Simhadri discloses that the trained object-detection algorithm generates an appropriately sized bounding box around the detected person. (Simhadri: ¶ [0037]).
Simhadri further expressly discloses that its bounding-box component employs a machine-learning or deep-learning algorithm to detect a person within a frame and generate a bounding box around each detected person in the frame. (Simhadri: Fig. 4; ¶ [0076]).
Begun in view of Simhadri are combinable because they are from the same field of endeavor of image processing (identifying characters in a video). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Simhadri’s model-based bounding-box detection into Begun’s content-recognition method. Begun and Simhadri concern analogous image and video-processing systems that detect people or faces in video frames and perform subsequent recognition related processing. Begun identifies characters based on facial or other biometric features detected from video frames, while Simhadri provides a known implementation for detecting persons and localizing each detected person within a respective bounding box.
A person of ordinary skill would have been motivated to use Simhadri’s trained object-detection model and bounding boxes in Begun’s system to localize and segment each detected person or face from the remainder of the sampled frame, thereby providing standardized regions for subsequent biometric-feature extraction and character identification. Such a modification would predictably improve localization of the content elements, reduce processing of irrelevant background portions of the frame, and facilitate Begun’s downstream feature matching and identification. The modification represents the use of a known model-based object localization technique for its established purpose in an analogous video-recognition system and would have yielded predictable results.
As to the newly added limitation:
“wherein identifying the detected content elements comprises, for each detected content element,”
Begun discloses detecting and identifying individual characters based on their respective detected biometric features. (Begun: Fig. 2, operations S210-S220; Fig. 5, operations S520-S530; ¶¶ [0058; 0061-0065]).
Simhadri discloses generating a respective bounding box around each detected person in a frame. Accordingly, the combined system performs the subsequent processing with respect to each detected person and its respective bounding box. (Simhadri: ¶ [0076]).
“performing alignment over each bounding box by searching for key points or landmarks over the bounding box”;
Simhadri discloses that a trained multi-pose estimation algorithm analyzes the portion of the frame within the bounding box and generates a heatmap showing keypoints or anatomical masks of the detected person, including primary joints and facial features such as the eyes, ears, nose, mouth, and chin. (Simhadri: ¶ [0038]).
Simhadri further discloses providing a plurality of keypoint predictions throughout the bounding box and selecting the highest predictions as the inferred locations of the relevant keypoints or anatomical masks. Simhadri therefore searches for keypoints or landmarks over the bounding box. (Simhadri: Fig. 7; ¶ [0085]).
“to obtain a geometric structure”;
Simhadri discloses that the resulting heatmap includes an arrangement of vertices corresponding to the detected keypoints or anatomical masks and lines connecting the vertices. The vertices are positioned over respective keypoints of the person within the bounding box, and coordinates define the locations of the keypoints. This arrangement of spatially positioned vertices and connecting lines constitutes a geometric structure of the detected content element. (Simhadri: ¶ [0038]; see also Fig. 3 and ¶¶ [0070-0071]).
Simhadri also discloses outputting coordinates or approximate coordinates for the detected keypoints, including the shoulders, hips, eyes, ears, and nose. These landmark coordinates further define the geometric structure of the detected person or face. (Simhadri: ¶ [0085]).
Simhadri additionally discloses that the bounding-box image may be substantially centered around the detected person and may be cropped, resized, and normalized to facilitate subsequent processing by the pose-estimation algorithm. Thus, Simhadri teaches transforming and normalizing the image region containing the detected content element. (Simhadri: ¶ [0088]).
However, Simhadri does not expressly disclose computing the alignment transformation from the landmark-defined geometric structure in the claimed manner.
Cao discloses obtaining alignment based on a landmark-defined geometric structure and a transformation, as recited “and obtaining alignment of the detected content element through one or more transformations including translation, scale, or rotation.”
Cao discloses face alignment by explicit shape regression in which facial landmarks are jointly regressed to infer a facial shape. Cao further discloses an alignment-estimation module that estimates the face shape for an input image. (Cao: Figs. 1 and 10; ¶¶ [0018-0019; 0022], and [0089-0093]).
Cao defines a face shape as a geometric structure comprising a plurality of facial landmarks, each represented by x- and y-coordinate values. (Cao: ¶[0053]).
Cao explains that image features can vary because of differences in scale or rotation. To provide invariance against those differences, Cao computes a similarity transform that normalizes a current landmark-defined shape to a mean shape. Cao further discloses estimating the mean shape by least-squares fitting of facial landmarks, including eye corners, a nose tip, a chin, and mouth corners. (Cao: ¶[0025]).
Cao therefore teaches obtaining alignment using a transformation computed from a geometric structure defined by facial landmarks. Cao expressly identifies scale and rotation as the variations addressed by the similarity transformation. Because claim 1 recites “one or more transformations” and lists translation, scale, or rotation disjunctively, Cao’s scale- and rotation-based similarity transformation satisfies the recited transformation requirement.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to further modify the combined Begun-Simhadri system to use Cao’s landmark-based similarity transformation to align each detected content element within its respective bounding box before Begun’s downstream biometric-feature matching and identification. In the resulting combination, Simhadri’s detection model would generate a bounding box around each detected person; Simhadri’s pose-estimation processing would search for keypoints throughout the bounding box and produce a geometric structure comprising landmark vertices, connecting lines, and landmark coordinates; and Cao’s similarity-transformation technique would use that landmark-defined geometric structure to normalize the bounded content element with respect to scale and/or rotation before the normalized content element is supplied to Begun’s feature-matching and character-identification processing.
The express motivation for the modification is provided by Cao, which explains that image features vary because of differences in scale and rotation and that a similarity transformation based on facial landmarks provides invariance against those variations. (Cao: ¶ [0025]). Applying Cao’s alignment technique to the bounded persons or faces of Simhadri would standardize the detected content elements before Begun’s biometric-feature extraction and database matching, thereby improving the consistency, robustness, and accuracy of the subsequent identification.
Accordingly, claim 1 is unpatentable over Begun in view of Simhadri and further in view of Cao.
Regarding Claim 2:
The proposed combination of Begun, Simhadri and Cao further discloses the method of claim 1, wherein identifying the detected content elements comprises:
for each detected content element;
performing alignment over each bounding box;
Simhadri discloses detecting persons within image data and generating bounding boxes corresponding to detected persons, wherein the region defined by each bounding box is cropped, resized, and formatted prior to further processing by downstream models (Simhadri: ¶[0088]). Such cropping, resizing, and normalization of a detected region constitutes alignment over each bounding box, as it standardizes the detected content element for subsequent analysis.
performing quality analysis over the aligned bounding boxes to generate analysis scores, each analysis score being associated with a detected content element; and
Simhadri further discloses that the detection process outputs a confidence score for each detected bounding box, the confidence score indicating the likelihood that the bounding box contains a valid person detection (Simhadri: ¶[0076]). The confidence score is generated per detected bounding box and therefore constitutes an analysis score associated with a detected content element, as recited.
performing matching on each detected content element associated with any analysis score meeting or exceeding a scoring threshold.
Simhadri teaches using the confidence score as a basis for determining whether a detected content element is accepted for further processing (Simhadri: ¶0076]), which corresponds to applying a scoring threshold.
Once the detected content element satisfies the scoring threshold, Begun teaches performing feature-based matching and classification of the detected content element, including calculating distances between extracted biometric feature vectors and grouping or classifying detected elements on a per-identify base (Begun: ¶¶[0098-0099,¶[0093]).
Thus, the combined disclose performing matching only on detected content elements whose associated analysis scores meet or exceed a threshold, as claimed.
Begun, Simhadri and Cao are combinable because they are from the same field of endeavor of image processing (identifying characters in a video). It would have been obvious to one of ordinary skill in the art before the effective filing date to combine Simhadri’s bounding box alignment and confidence score thresholding with Begun’s feature based matching and classification techniques. The suggestion/motivation for doing so is to improve recognition accuracy and computational efficiency by filtering low-quality detections before performing matching. Accordingly, Claim 2 is unpatentable over Begun, Simhadri and Cao.
Regarding Claim 3:
The proposed combination of Begun, Simhadri and Cao further discloses the method of claim 2, wherein performing matching on each detected content element comprises:
extracting embedding associated with the content element;
Begun expressly teaches generating feature embeddings (biometric feature vectors) from detected content elements for use in downstream analysis and classification (Begun: ¶[0093]), describing use of embedding based techniques such as t-SNE for representing biometric features).
matching the extracted embedding against stored embeddings to locate an identity of the element; and
Begun further discloses calculating distances between extracted biometric feature embeddings and other stored biometric features in order classify, group, or associate detected elements on a per person (identity) basis (Begun: ¶¶[0098-0099]). Such distance-based comparison of embeddings against a stored embeddings constitutes matching the extracted embedding to locate an identity.
outputting the located identity as an identity of the content element.
Begun teaches assigning detected biometric features to an identified person or group and using that classification as the recognized identity of the detected content element (Begun: ¶[0099]), which corresponds to outputting the located identity as claimed.
Accordingly, Claim 3 is unpatentable over Begun, Simhadri and Cao.
Regarding Claim 4:
The proposed combination of Begun, Simhadri and Cao discloses the method of claim 3, wherein matching the extracted embedding against the stored embeddings to locate an identity of the content element is performed on a server.
Simhadri discloses wherein matching the extracted embedding against the stored embeddings to locate an identity of the content element is performed on a server.
Begun discloses matching detected facial content features against stored reference data to identify a content element. Specifically Begun teaches a Face Recognition Engine (Begun: Fig. 2 S210 ‘Face Detection Engine’) that compares detected face information against a ‘faces database’ (Begun: Fig. 2; ¶[0058]) to determine an identity of the detected content element (e.g., actor identification). See Fig. 2 (face detection engine – face recognition engine – faces database – identified actors), and ¶¶[0098-0099], which describe calculating distances between biometric features and classify and grouping features on a person basis to identify individuals.
Simhadri compliments Begun by teaching a recognition pipeline in which detected and qualified content elements are processed by downstream recognition functionality once they satisfy confidence thresholds, reinforcing that recognition matching is a discrete processing stage suitable for modular deployment; (Simhadri: ¶¶[0076-0088]).
Simhadri expressly teaches performing processing in a server-based architecture, wherein computationally intensive operations and reference data are maintained and executed on a server rather than exclusively on a client device. See Simhadri, system architecture discussion describing server-side processing in communication with client devices (e.g., server performing core processing used stored data); (Simhadri: ¶[0163]; ¶[0154]; ¶[0169]; ¶¶[0173-0175]; Fig. 25).
Begun, Simhadri and Cao are combinable because they are from the same field of endeavor of image processing (e.g., identifying characters in a video).
It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to perform the matching of extracted embeddings against stored embeddings in Begun on a server, as taught by Simhadri. The suggestion/motivation for doing so is that centralizing the matching process on a server improves computational efficiency, enables shared access to large reference libraries, and facilitates maintenance and updating of stored embeddings, all of which improve design considerations in content recognition. Accordingly, it would have been obvious to combine Begun, Simhadri and Cao to arrive at the subject matter of claim 4.
Regarding Claim 5:
The proposed combination of Begun, Simhadri and Cao further discloses the method of claim 1, further comprising:
for each of the identified content elements:
searching for at least one matching work associated with the identified content element;
Begun explains identifying characters in a sampled video and using those identified characters to identify the video content by matching against a content-specific character database/character list (e.g., locating video content works associated with the identified character(s)); (Begun: ¶[0007]).
Begun further describes operations in which a first character is identified and used for identifying the video content (with the identification relying on stored character information for video content); (Begun: ¶¶[0061-0066]; ¶¶[0070-0073]; Fig. 6).
grouping the at least one matching work with the identified content element into a set of works associated with the identified content element; and
Begun’s approach of identifying video content by matching identified character(s) to stored character lists for respective video contents necessarily yields, for a given identified character, a collection of candidate video content items in the database associated with that character (i.e., a “set” of works associated with the identified content element, as claimed); (Begun: ¶[0007], ¶¶[0070-0075]; Fig. 6).
determining whether an intersecting work exists between the sets of works.
Simhadri expressly discloses evaluating overlap between result sets using an intersection operation (e.g., intersection over union and set based overlap determinations), which teaches determining whether an intersecting item exists between multiple sets derived from detected elements; (Simhadri: ¶[0082]).
Begun, Simhadri and Cao are combinable because they are from the same field of endeavor of image processing (identifying characters in a video). It would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate Simhadri’s explicit set intersection determination into Begun’s multi-element candidate work aggregation framework. The suggestion/motivation for doing so is to resolve a single source work common to multiple detected content elements and to improve the identification accuracy and reduce ambiguity when multiple element specific candidate sets are present. Accordingly, Claim 5 is unpatentable over Begun, Simhadri and Cao.
Regarding Claim 6:
The proposed combination of Begun, Simhadri and Cao further discloses the method of claim 5, further comprising:
finding one intersecting work between the sets of works, and outputting the intersecting work as an identity of the source content.
Begun further discloses that, after comparing the per-element lists of candidate works and identifying common candidate content, the system proceeds with content identification/output based on the identified matching content. In particular, Begun describes identifying the video content that includes multiple identified characters after comparing the lists associated with those characters (i.e., narrowing to the common works) and thereby determining the identified content; (Begun: Fig. 2 ‘bottom left wherein it is determined from all the identified actors that the video frames are from Ocean’s Eleven’; ¶¶[0068-0069] ‘comparing lists corresponding to each character and identifying the video content including two or more characters), and (Begun: ¶[0070] ‘continuing the method flow after the identification step at S540’).
This corresponds to claim 6’s requirement of finding one intersecting work (i.e., a single common work after comparison) and outputting the intersecting work as an identity of the source content, because the common result of the per-element lists is used as the identified content output; (Begun ¶¶[0068-0069]).
Begun explicitly discloses identifying (and thereby outputting/returning) the video content (the source content) based on matching identified characters to stored character lists for video contents in a database, e.g., the workflow resolves to the video content work that satisfies the identified character matching constraints; (Begun: ¶0007], ¶¶[0070-0071]).
Accordingly, Claim 6 is unpatentable over Begun, Simhadri and Cao.
Regarding Claim 8:
The proposed combination of Begun, Simhadri and Cao further discloses the method of claim 5, further comprising:
finding more than one intersecting work among the sets of works:
Begun’s approach of identifying video content by matching identified character(s) to stored character lists for respective video contents necessarily yields, for a given identified character, a collection of candidate video content items in the database associated with that character (i.e., a “set” of works associated with the identified content element, as claimed); (Begun: ¶[0007], ¶¶[0070-0075]; Fig. 6).
Simhadri expressly discloses evaluating overlap between result sets using an intersection operation (e.g., intersection over union and set based overlap determinations), which teaches determining whether an intersecting item exists between multiple sets derived from detected elements; (Simhadri: ¶[0082]).
detecting an additional content element from the sampled source content;
Simhadri describes ML based detection of people and objects in frames and continuing processing based on detections (object and person detection context). See, e.g., Simhadri ¶[0082] (YOLO based detector), which supports the art recognized approach of detecting additional elements as needed from visual content.
identifying the detected additional content element;
Begun describes identifying corresponding characters based on detected biometric features, i.e., identification of detected persons and characters. See Begun ¶[0062] (identifying corresponding characters from biometric features) and ¶[0063] (repeat the same character- based process on another frame).
searching for at least one additional matching work associated with the identified additional content element;
Begun uses identified characters (e.g., actors and actresses) with content databases (e.g., IMDB) to identify video content based on identified elements. See Begun ¶[0061] (use content specific character database to identify video content) and ¶[0062] (explicitly referencing IMDB type sources in the identification workflow).
grouping the at least one additional matching work associated with the identified additional content element into an additional set of works; and
Begun teaches maintaining lists/databases of identified and non-identified biometric features and character lists associated with content identification workflows (i.e., aggregating candidate identity evidence for later matching, including using a Videos List/Short List during detection and identification.
determining whether an intersecting work exists between the sets of works and the additional set of works.
Begun teaches that identification is performed “based on” matching identified characters against character lists for video content, and when not resolved repeats on another frame to reach identification (i.e., iterative narrowing to a resolvable identification).
Begun, Simhadri and Cao are combinable because they are from the same field of endeavor of image processing (identifying characters in a video). It would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate Simhadri’s ML based detection of people and objects in frames into Begun’s multi-element candidate work aggregation framework. The suggestion/motivation for doing so is to resolve a single source work common to multiple detected content elements and to improve the identification accuracy and reduce ambiguity when multiple element specific candidate sets are present. Accordingly, Claim 8 is unpatentable over Begun, Simhadri and Cao.
Regarding Claim 9:
The proposed combination of Begun, Simhadri and Cao further discloses the method of claim 8, further comprising:
finding one intersecting work among the sets of works and the additional set of works, and outputting the intersecting work as an identity of the source content.
As disclosed in the rejection of Claim 8 the sets of works and the additional set of works are narrowed down to a Singular Video, from a List of Videos (Begun: ¶[0007], ¶¶[0070-0075]; Fig. 6 and bottom left of Fig. 24).
Accordingly, Claim 9 is unpatentable over Begun, Simhadri and Cao.
Regarding Claim 11:
The proposed combination of Begun, Simhadri and Cao further discloses the method of claim 1, wherein the source content comprises an audio, visual, or audiovisual work.
Begun teaches wherein “A frame may include not only image data but also acoustic data.” (Begun: Fig. 6 Steps S510, S610 and S620 ¶[0071]).
Accordingly, Claim 11 is unpatentable over Begun, Simhadri and Cao.
10. Claims 7 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Begun, Simhadri and Cao as applied to claims 5 and 8 above, and further in view of Shen et al. (US 2021/0117691), hereinafter ‘Shen’.
Regarding Claim 7:
The proposed combination of Begun, Simhadri and Cao further discloses the method of claim 5, comprising:
finding no intersecting work between the sets of works,
Begun explains that when the system fails to identify the content element based on the current frame sample, the identification attempt is unsuccessful and the process does not resolve to a determinative identity at that stage (Begun: ¶[0113]). This corresponds to the claimed condition where the current candidate resolution attempt yields no successful intersection of results.
Note that Simhadri expressly uses “intersection” terminology (intersection over union) in connection with bounding boxes and object detection; (Simhadri: ¶[0082]).
restarting content sampling process.
Begun expressly teaches performing the identification process on another frame when identification is not achieved, thereby restarting the sampling and analysis process with new content; (Begun: ¶[0113]).
Simhadri also discloses accessing and operating on video frames as discrete samples, supporting repeated sampling from source content after an unsuccessful attempt; (Simhadri: ¶¶[0059-0060]).
Begun, Simhadri and Cao do not expressly disclose discarding the sets of works.
Shen discloses discarding the sets of works.
Shen teaches discarding frames or samples that are not analyzed or not useful for continued processing, i.e., abandoning a non-productive attempt rather than carrying it forward. In the context of claim 5’s grouping of candidate works, this teaches discarding the current sets when they do no resolve the identity; (Shen ¶[0032]).
Begun, Simhadri, Cao & Shen are combinable because they are from the same field of endeavor of image processing (e.g., identifying characters in a video). It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to discard a non-resolving candidate attempt and restart sampling with new content, as taught by Begun and Shen, using routine frame-based processing as taught by Simhadri. The suggestion/motivation for doing so is to improve robustness and efficiency when the current attempt fails to yield a determinative identification. Accordingly, it would have been obvious to combine Begun, Simhadri, Cao, and Shen to arrive at the subject matter of claim 7.
Regarding Claim 10:
The proposed combination of Begun, Simhadri and Cao further discloses the method of claim 8, further comprising:
finding no intersecting work between the sets of works,
Begun explains that when the system fails to identify the content element based on the current frame sample, the identification attempt is unsuccessful and the process does not resolve to a determinative identity at that stage (Begun: ¶[0113]). This corresponds to the claimed condition where the current candidate resolution attempt yields no successful intersection of results.
Note that Simhadri expressly uses “intersection” terminology (intersection over union) in connection with bounding boxes and object detection; (Simhadri: ¶[0082]).
restarting content sampling process.
Begun expressly teaches performing the identification process on another frame when identification is not achieved, thereby restarting the sampling and analysis process with new content; (Begun: ¶[0113]).
Simhadri also discloses accessing and operating on video frames as discrete samples, supporting repeated sampling from source content after an unsuccessful attempt; (Simhadri: ¶¶[0059-0060]).
Begun, Simhadri and Cao do not expressly disclose discarding the sets of works.
Shen discloses discarding the sets of works.
Shen teaches discarding frames or samples that are not analyzed or not useful for continued processing, i.e., abandoning a non-productive attempt rather than carrying it forward. In the context of claim 8’s grouping of candidate works, this teaches discarding the current sets when they do no resolve the identity; (Shen ¶[0032]).
Begun, Simhadri, Cao & Shen are combinable because they are from the same field of endeavor of image processing (e.g., identifying characters in a video). It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to discard a non-resolving candidate attempt and restart sampling with new content, as taught by Begun and Shen, using routine frame-based processing as taught by Simhadri. The suggestion/motivation for doing so is to improve robustness and efficiency when the current attempt fails to yield a determinative identification. Accordingly, it would have been obvious to combine Begun, Simhadri, Cao, and Shen to arrive at the subject matter of claim 10.
11. Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Begun, Simhadri and Cao as applied to claim 1 above, and further in view of Zadeh et al. (US 11,914,674), hereinafter ‘Zadeh’.
Regarding Claim 12:
The proposed combination of Begun, Simhadri and Cao further discloses the method of claim 1, wherein the content element is a person, vehicle,
Begun teaches detecting biometric features (e.g., facial/biometric features) from the extracted frame. These biometric feature-bearing regions correspond to content elements (e.g., persons/faces) detected from the sample source content; (Begun: Fig. 5 flowchart: ‘detect biometric features from extracted frame’ at S530; ¶[0062]).
Begun expressly discloses sampling audiovisual source content and performing object detection on sampled frames (Begun: ¶¶[0022-0028], and ¶¶[0031-0036] and Fig. 2)
Begun further discloses wherein the biometric features include ‘faces, voices, gaits, ears, and hand shapes’ Fig. 1).
Simhadri describes ML based detection of people and objects in frames and continuing processing based on detections (object and person detection context). See, e.g., Simhadri ¶¶[0078-0082] (YOLO based detector), which supports the art recognized approach of detecting additional elements as needed from visual content, such as logos (¶0051]).
Note that it is well-known class list for a standard YOLO dataset to include ‘person, vehicle, plant, animals, clothing, sign (such as traffic light or stop sign), textual or numeric information, landmark locations, posters (which correspond to visual works).
Begun, Simhadri and Cao do not expressly disclose wherein the content element is a building, city, geographic feature, slogan, symbol, word, jingle, brand, trade name, trademark.
Zadeh discloses wherein the content element is a building, city, geographic feature, slogan, symbol, word, jingle, brand, trade name, trademark.
Zadeh discloses intelligent recognition of names, patterns, and semantic constructs, which correspond to brand names (Zadeh: Col. 109, lines 49-67), trade names (Zadeh: Col. 109, lines 49-67), trademarks (Zadeh: Col. 109, lines 49-67), slogans (Zaheh: Col. 105, lines 1-14), words (Zadeh: Col. 261 lines 62 – Col. 262, line 20), symbol (Zadeh: Col. 201, lines 37-49) and jingles (Zadeh: Col. 105, lines 1-14).
Zadeh discloses recognition of structured semantic entities and contextual constructs, which encompass buildings (Zadeh: Col. 14, lines 2-26), cities (Zadeh: Col. 212 line 56 – Col. 213, line 4), geographic entities (Zadeh: Col. 200, lines 19-44), landmarks (Zadeh: Col. 14, lines 2-26) and location-based identifiers (Zadeh: Col. 14, lines 2-26).
Begun, Simhadri, Cao & Zadeh are combinable because they are from the same field of endeavor of image processing (e.g., identifying objects in a video). It would have been obvious to one of ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate Zadeh’s semantic name and pattern recognition capabilities into the object detection and content recognition framework of Begun and Simhadri. The suggestion/motivation for doing so is to broaden the range of detectable content elements beyond purely visual object classes to include semantic identifiers such as brand names, trademarks, geographic entities, and structured concepts to improve robustness and coverage of content element identification across audiovisual content. Accordingly, it would have been obvious to combine Begun, Simhadri, Cao, and Zadeh to arrive at the subject matter of claim 12.
Conclusion
12. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
13. Any inquiry concerning this communication or earlier communications from the examiner should be directed to NEIL R MCLEAN whose telephone number is (571)270-1679. The examiner can normally be reached Monday-Thursday, 6AM - 4PM, PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Akwasi M Sarpong can be reached at 571.270.3438. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NEIL R MCLEAN/Primary Examiner, Art Unit 2681