Prosecution Insights
Last updated: August 15, 2026
Application No. 18/804,573

TTML PLACEMENT INFLUENCED BY OBJECT DETECTION

Final Rejection §103
Filed
Aug 14, 2024
Examiner
LIN, JASON K
Art Unit
2425
Tech Center
2400 — Computer Networks
Assignee
Comcast Cable Communications LLC
OA Round
4 (Final)
49%
Grant Probability
Moderate
5-6
OA Rounds
1y 8m
Est. Remaining
83%
With Interview

Examiner Intelligence

Grants 49% of resolved cases
49%
Career Allowance Rate
225 granted / 460 resolved
-9.1% vs TC avg
Strong +34% interview lift
Without
With
+33.9%
Interview Lift
resolved cases with interview
Typical timeline
3y 8m
Avg Prosecution
22 currently pending
Career history
488
Total Applications
across all art units

Statute-Specific Performance

§101
5.6%
-34.4% vs TC avg
§103
63.4%
+23.4% vs TC avg
§102
14.5%
-25.5% vs TC avg
§112
8.8%
-31.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 460 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION This office action is responsive to application No. 18/804,573 filed on 02/04/2026. Claim(s) 1-20 is/are pending and have been examined. Claim Objections Claim(s) 17 is/are objected to because of the following informalities: Claim 17 recites: “preventing output, in the video content and for the first time period, of an overlay object in the one or more regions during the first time period” an overlay object has already been recited previously in the first two paragraphs of claim 17: “receiving, by a computing device, an overlay object…” “…one or more first content objects to remain at least partially unobscured by the overlay object” Additionally Claim 1 recites similar limitation(s) and recites: “preventing output, in the video content and for the first time period, of the overlay object in the one or more regions” If Applicant intends “overlay object” to refer to the same overlay object, please amend to: -- preventing output, in the video content and for the first time period, of the overlay object in the one or more regions -- Appropriate correction is required. Response to Arguments Applicant's arguments filed 02/04/2026 have been fully considered but they are not persuasive. Applicants assert that “ None of the cited references teach, disclose, or otherwise suggest at least, as recited by claim 1 (as amended), "determining, based on a first importance value corresponding to the one or more second content objects, a first time period during which the one or more second content objects should not be obstructed by one or more overlay objects;" ''preventing output, in the video content and for the first time period, of the overlay object in the one or more regions;" and "after determining that the first time period has elapsed, causing output of the video content with the overlay object inserted into the one or more regions." In response, the Examiner respectfully disagrees. Note that including, but not limited to the paragraphs cited in Zhang, we see that in a sports video category moving objects may be given a weight value of one. Detected features, which include movement, may be situated at different locations within different frames that occupy a certain duration of the input video. With a weight of one, having a high level of importance, these identified features can be locations that should not, if possible, be occluded by overlaid content. As features may be situated at different locations within frames that occupy a certain duration of the video, motion of the object may only be for a certain duration (first time period), where after that duration, there may be no more motion of the object. When that happens, a weight of one which is given to a moving object, would no longer carry a weight of 1. Thus, lowering the importance of the corresponding one or more regions, allowing for overlay object to be inserted after the first time period has elapsed. Therefore, based on the above and the Office Action below. The combination of Zhang and Gupta continue to teach the claimed limitation(s). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-8, 10-12, 14-17, and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US 2022/0368979) in view of Gupta et al. (US 2019/0349626). Consider claim 1, Zhang teaches a method comprising: receiving, by a computing device, an overlay object for insertion into video content (Paragraph 0037 teaches videos that are streamed to a user can include additional content, e.g., provided by a content provider 108, that is overlaid on top of the original video stream); identifying, based on one or more first content objects identified in one or more frames of a plurality of frames of the video content, one or more regions of the one or more frames, where action of the one or more second content objects is to occur (Paragraph 0003 teaches determining a location at which to position overlaid content during video display. Paragraph 0007 teaches determining location at which to position overlaid content during video display based on duration or number of frames over which the overlaid content is to be provided within the video. Paragraph 0028 teaches assigning different levels of importance to each detected feature type. For sports-related video a high importance may be assigned for not overlaying portions of the video where there is motion. Paragraph 0029 teaches based on the adjusted confidence values for the different feature types, the video system can identify one or more locations at which the overlay content can be displayed. Paragraph 0088 teaches determining based on the aggregated and adjusted confidence scores, a location at which to display the overlaid content during display of the video. Overlay location is determined based on desired display duration. Paragraph 0060 teaches other types of important features that a user may prefer to not be occluded can be moving objects. For example, objects that are in motion are generally more likely to convey important content to the viewer than static objects and are therefore generally less suitable to be occluded by overlaid content. The moving object detector 310e can detect moving objects by detecting movement between adjacent sampled frames, based on color-space differences between frames. Processing can be performed for each frame, other than first and last frames, using a previous, current, and next frame, with calculations producing a confidence score for the current frame. Paragraph 0067 teaches each feature type weight for a category can be, for example, a value between zero and one. A weight value of less than one, e.g., 0.5, for a feature type for a video category can indicate that the feature type is less important for the video category, a sports video category may have a weight of one for moving object features, a weight of one for text features, e.g., for scoreboards or statistics, and a weight of 0.8 for human face features. Paragraph 0073 teaches detected features, as represented by aggregated and adjusted confidence scores, may be situated at different locations within different frames that occupy a certain duration of the input video 302. Locations that include identified features can be locations that should not, if possible, be occluded by overlaid content, e.g., to ensure that important content of the underlying input video 302 is not occluded. As another example, an overlaid content location 320a can be a location which corresponds to confidence scores that are below a predetermined threshold. As locations with motion are taken into consideration when selecting a location for the overlay object, successive sequential frames are compared to identify motion, where regions with motion, can correspond to a second region that is identified. These regions are granted a higher importance/weight, which should not be occluded by overlaid content); determining, based on a first importance value corresponding to the one or more second content objects, a first time period during which the one or more second content objects should not be obstructed by one or more overlay objects (Paragraph 0028 teaches assigning different levels of importance to each detected feature type. For sports-related video a high importance may be assigned for not overlaying portions of the video where there is motion. Paragraph 0060 teaches other types of important features that a user may prefer to not be occluded can be moving objects. For example, objects that are in motion are generally more likely to convey important content to the viewer than static objects and are therefore generally less suitable to be occluded by overlaid content. The moving object detector 310e can detect moving objects by detecting movement between adjacent sampled frames, based on color-space differences between frames. Processing can be performed for each frame, other than first and last frames, using a previous, current, and next frame, with calculations producing a confidence score for the current frame. Paragraph 0067 teaches each feature type weight for a category can be, for example, a value between zero and one. A weight value of less than one, e.g., 0.5, for a feature type for a video category can indicate that the feature type is less important for the video category, a sports video category may have a weight of one for moving object features, a weight of one for text features, e.g., for scoreboards or statistics, and a weight of 0.8 for human face features. Paragraph 0073 teaches detected features, as represented by aggregated and adjusted confidence scores, may be situated at different locations within different frames that occupy a certain duration of the input video 302. Locations that include identified features can be locations that should not, if possible, be occluded by overlaid content, e.g., to ensure that important content of the underlying input video 302 is not occluded. For example, if the video category is sports, a weight value of 1 may be given for moving object features, where moving objects may not be occluded. Movement of moving objects contain motion between adjacent frames, in different locations that occupy a certain duration of the input video. Thus, during this duration of motion of the object, the object must not be obstructed by an overlay object); preventing output, in the video content and for the first time period, of the overlay object in the one or more regions; and after determining that the first time period has elapsed, causing, output of the video content with the overlay object inserted into the one or more regions (Paragraph 0003 teaches providing the overlaid content for display at the determined location in the video. Paragraph 0028, 0060 teaches for sports related video, a high importance may be assigned for not overlaying portions of the video where there is motion. Paragraph 0067 teaches sports category may have a weight of one for moving object features. Paragraph 0089 teaches video processing system 110 provides the overlaid content for display at the determined location in the video. Video processing system 110 can provide an overlaid content location 320a and a corresponding time offset 320 for an overlaid content item, and in some implementations, the overlaid content item itself, to the content platform 106. Paragraph 0090 teaches video processing system 110 and/or content platform 106 can display an overlaid content item on top of video content without occluding important video content. Paragraph 0071 teaches one or more existing content slots used to display overlay content on the input video 302. Paragraph 0073 teaches detected features, as represented by aggregated and adjusted confidence scores, may be situated at different locations within different frames that occupy a certain duration of the input video 302. An overlaid content location 320a can be determined so overlaid content that is displayed starting at a time offset 320b and for a specified duration of time that includes multiple frames is not positioned anywhere that the detected features are situated across the multiple frames. An identified overlaid content location 320a can be a location generally or substantially outside locations that include the identified features in a sequence of video frames corresponding to the input duration. Locations that include identified features can be locations that should not, if possible, be occluded by overlaid content, e.g., to ensure that important content of the underlying input video 302 is not occluded. As another example, an overlaid content location 320a can be a location which corresponds to confidence scores that are below a predetermined threshold. Paragraph 0078 teaches the overlay location identifier 318 can determine location(s) that correspond to low confidence scores that persist across a set of frames, for a location for display of overlay content for a set of frames equaling a desired duration. Paragraph 0080 teaches a content slot 506, which is located outside of colored areas corresponding to high confidence scores, corresponds to a recommended location for display of an overlaid content item for a specified duration starting at, or at least including, a time offset of the talk show video corresponding to the video frame 200. Based on importance, overlays may be prevented from being output over areas with important content, such as moving objects. Movement of moving objects contain motion between adjacent frames, in different locations that occupy a certain duration of the input video. Thus, during this duration of motion of the object, the object must not be obstructed by an overlay object. However, past the duration of the motion of the object, where the corresponding areas may no longer have movement, those areas become unimportant now as importance was determined based on motion of object(s). For areas having unimportant or less important content, areas where confidence score are below a threshold, outside of areas of high confidence scores, overlay content may be displayed. Thus, during times where objects have motion, overlay is prevented from display on those corresponding areas, but when there is no longer motion after a certain duration, those same areas become unimportant, and overlays may be displayed on those corresponding areas). Zhang does not explicitly teach where action is to occur, is where future action is of one or more second content objects is predicted to occur; In an analogous art, Gupta teaches identifying, based on one or more first content objects, one or regions of the one or more frames, where future action of one or more second content objects is predicted to occur (Abstract, Figs.1, Paragraph 0027-0028, 0034-0035, 0041, 0043; Fig.6, Paragraph 0107-0110). Therefore, it would have been obvious to a person of ordinary skill in the art to modify the system of Zhang to include identifying, based on one or more first content objects, one or regions of the one or more frames, where future action of one or more second content objects is predicted to occur, as taught by Gupta, for the advantage of generating for display an overlay in a location that does not overlap with any of the first location, the second location, and the projected location (Gupta – Abstract, Paragraph 0110), better ensuring that viewers would not suffer a loss of enjoyment when consuming media content by avoiding overlays obscuring items of importance. Consider claim 11, Zhang teaches a method comprising: determining, by a computing device, a plurality of content objects present in at least a portion of a plurality of frames of video content (Paragraph 0027; Fig.3, Paragraph 0051-0060); determining, for each of the plurality of content objects, an importance value that indicates a priority of the content object remaining unobscured; identifying one or more first content objects, of the plurality of content objects, associated with highest importance values (Paragraph 0028-0029, 0067-0068; Paragraph 0028 teaches assigning different levels of importance to each detected feature type. A high importance/weight may be assigned for not overlaying portion of the video that include human faces. For sports-related video a high importance may be assigned for not overlaying portions of the video where there is motion. Paragraph 0029 teaches based on the adjusted confidence values for the different feature types, the video system can identify one or more locations at which the overlay content can be displayed. Paragraph 0060 teaches other types of important features that a user may prefer to not be occluded can be moving objects. For example, objects that are in motion are generally more likely to convey important content to the viewer than static objects and are therefore generally less suitable to be occluded by overlaid content. The moving object detector 310e can detect moving objects by detecting movement between adjacent sampled frames, based on color-space differences between frames. Processing can be performed for each frame, other than first and last frames, using a previous, current, and next frame, with calculations producing a confidence score for the current frame. Paragraph 0067 teaches each feature type weight for a category can be, for example, a value between zero and one. A weight value of less than one, e.g., 0.5, for a feature type for a video category can indicate that the feature type is less important for the video category, a sports video category may have a weight of one for moving object features, a weight of one for text features, e.g., for scoreboards or statistics, and a weight of 0.8 for human face features. Paragraph 0073 teaches detected features, as represented by aggregated and adjusted confidence scores, may be situated at different locations within different frames that occupy a certain duration of the input video 302. Locations that include identified features can be locations that should not, if possible, be occluded by overlaid content, e.g., to ensure that important content of the underlying input video 302 is not occluded. As another example, an overlaid content location 320a can be a location which corresponds to confidence scores that are below a predetermined threshold. Fig.5, Paragraph 0079 teaches a confidence score visualization 500. A content slot 506. Paragraph 0080 teaches a content slot 506, located outside of the colored areas corresponding to high confidence scores, that corresponds to a recommended location for display of an overlaid content item for a specified duration. Paragraph 0081 teaches although a single frame and single content slot visualization are shown, multiple frame visualizations, e.g., as a collection of still images or as a modified video stream, can be generated. There may be different feature types at different region(s) of the frame(s) of the video, where one of these feature types may be a face/person. A face/person may be a content object that is located at a particular first region of frame(s) such as 502a-Fig.5. These regions are granted a higher importance/weight, which should not be occluded by overlaid content. Additionally in a sports video, moving objects, scoreboards, etc may be assigned a weight value of 1); identifying, based on one or more first content objects identified in one or more regions of the one or more frames, where action of the one or more second content objects, of the plurality of content objects is to occur (Paragraph 0003 teaches determining a location at which to position overlaid content during video display. Paragraph 0007 teaches determining location at which to position overlaid content during video display based on duration or number of frames over which the overlaid content is to be provided within the video. Paragraph 0028 teaches assigning different levels of importance to each detected feature type. For sports-related video a high importance may be assigned for not overlaying portions of the video where there is motion. Paragraph 0029 teaches based on the adjusted confidence values for the different feature types, the video system can identify one or more locations at which the overlay content can be displayed. Paragraph 0088 teaches determining based on the aggregated and adjusted confidence scores, a location at which to display the overlaid content during display of the video. Overlay location is determined based on desired display duration. Paragraph 0060 teaches other types of important features that a user may prefer to not be occluded can be moving objects. For example, objects that are in motion are generally more likely to convey important content to the viewer than static objects and are therefore generally less suitable to be occluded by overlaid content. The moving object detector 310e can detect moving objects by detecting movement between adjacent sampled frames, based on color-space differences between frames. Processing can be performed for each frame, other than first and last frames, using a previous, current, and next frame, with calculations producing a confidence score for the current frame. Paragraph 0067 teaches each feature type weight for a category can be, for example, a value between zero and one. A weight value of less than one, e.g., 0.5, for a feature type for a video category can indicate that the feature type is less important for the video category, a sports video category may have a weight of one for moving object features, a weight of one for text features, e.g., for scoreboards or statistics, and a weight of 0.8 for human face features. Paragraph 0073 teaches detected features, as represented by aggregated and adjusted confidence scores, may be situated at different locations within different frames that occupy a certain duration of the input video 302. Locations that include identified features can be locations that should not, if possible, be occluded by overlaid content, e.g., to ensure that important content of the underlying input video 302 is not occluded. As another example, an overlaid content location 320a can be a location which corresponds to confidence scores that are below a predetermined threshold. As locations with motion are taken into consideration when selecting a location for the overlay object, successive sequential frames are compared to identify motion, where regions with motion, can correspond to a second region that is identified. These regions are granted a higher importance/weight, which should not be occluded by overlaid content); determining, based on a first importance value corresponding to the one or more second content objects, a first time period during which the one or more second content objects should not be obstructed by one or more overlay objects (Paragraph 0028 teaches assigning different levels of importance to each detected feature type. For sports-related video a high importance may be assigned for not overlaying portions of the video where there is motion. Paragraph 0060 teaches other types of important features that a user may prefer to not be occluded can be moving objects. For example, objects that are in motion are generally more likely to convey important content to the viewer than static objects and are therefore generally less suitable to be occluded by overlaid content. The moving object detector 310e can detect moving objects by detecting movement between adjacent sampled frames, based on color-space differences between frames. Processing can be performed for each frame, other than first and last frames, using a previous, current, and next frame, with calculations producing a confidence score for the current frame. Paragraph 0067 teaches each feature type weight for a category can be, for example, a value between zero and one. A weight value of less than one, e.g., 0.5, for a feature type for a video category can indicate that the feature type is less important for the video category, a sports video category may have a weight of one for moving object features, a weight of one for text features, e.g., for scoreboards or statistics, and a weight of 0.8 for human face features. Paragraph 0073 teaches detected features, as represented by aggregated and adjusted confidence scores, may be situated at different locations within different frames that occupy a certain duration of the input video 302. Locations that include identified features can be locations that should not, if possible, be occluded by overlaid content, e.g., to ensure that important content of the underlying input video 302 is not occluded. For example, if the video category is sports, a weight value of 1 may be given for moving object features, where moving objects may not be occluded. Movement of moving objects contain motion between adjacent frames, in different locations that occupy a certain duration of the input video. Thus, during this duration of motion of the object, the object must not be obstructed by an overlay object); preventing output, in the video content and for the first time period, of an overlay object in the one or more regions; and after determining that the first time period has elapsed, causing, output of the video content with the overlay object inserted into the one or more regions (Paragraph 0003 teaches providing the overlaid content for display at the determined location in the video. Paragraph 0028, 0060 teaches for sports related video, a high importance may be assigned for not overlaying portions of the video where there is motion. Paragraph 0067 teaches sports category may have a weight of one for moving object features. Paragraph 0089 teaches video processing system 110 provides the overlaid content for display at the determined location in the video. Video processing system 110 can provide an overlaid content location 320a and a corresponding time offset 320 for an overlaid content item, and in some implementations, the overlaid content item itself, to the content platform 106. Paragraph 0090 teaches video processing system 110 and/or content platform 106 can display an overlaid content item on top of video content without occluding important video content. Paragraph 0071 teaches one or more existing content slots used to display overlay content on the input video 302. Paragraph 0073 teaches detected features, as represented by aggregated and adjusted confidence scores, may be situated at different locations within different frames that occupy a certain duration of the input video 302. An overlaid content location 320a can be determined so overlaid content that is displayed starting at a time offset 320b and for a specified duration of time that includes multiple frames is not positioned anywhere that the detected features are situated across the multiple frames. An identified overlaid content location 320a can be a location generally or substantially outside locations that include the identified features in a sequence of video frames corresponding to the input duration. Locations that include identified features can be locations that should not, if possible, be occluded by overlaid content, e.g., to ensure that important content of the underlying input video 302 is not occluded. As another example, an overlaid content location 320a can be a location which corresponds to confidence scores that are below a predetermined threshold. Paragraph 0078 teaches the overlay location identifier 318 can determine location(s) that correspond to low confidence scores that persist across a set of frames, for a location for display of overlay content for a set of frames equaling a desired duration. Paragraph 0080 teaches a content slot 506, which is located outside of colored areas corresponding to high confidence scores, corresponds to a recommended location for display of an overlaid content item for a specified duration starting at, or at least including, a time offset of the talk show video corresponding to the video frame 200. Based on importance, overlays may be prevented from being output over areas with important content, such as moving objects. Movement of moving objects contain motion between adjacent frames, in different locations that occupy a certain duration of the input video. Thus, during this duration of motion of the object, the object must not be obstructed by an overlay object. However, past the duration of the motion of the object, where the corresponding areas may no longer have movement, those areas become unimportant now as importance was determined based on motion of object(s). For areas having unimportant or less important content, areas where confidence score are below a threshold, outside of areas of high confidence scores, overlay content may be displayed. Thus, during times where objects have motion, overlay is prevented from display on those corresponding areas, but when there is no longer motion after a certain duration, those same areas become unimportant, and overlays may be displayed on those corresponding areas). Zhang does not explicitly teach where action is to occur, is where future action of one or more second content objects, of the plurality of content objects, is predicted to occur; In an analogous art, Gupta teaches identifying, based on the one or more first content objects, one or more regions of the one or more frames, where future action of the one or more second content objects, of the plurality of content objects, is predicted to occur (Abstract, Figs.1, Paragraph 0027-0028, 0034-0035, 0041, 0043; Fig.6, Paragraph 0107-0110). Therefore, it would have been obvious to a person of ordinary skill in the art to modify the system of Zhang to include identifying, based on the one or more first content objects, one or more regions of the one or more frames, where future action of the one or more second content objects, of the plurality of content objects, is predicted to occur, as taught by Gupta, for the advantage of generating for display an overlay in a location that does not overlap with any of the first location, the second location, and the projected location (Gupta – Abstract, Paragraph 0110), better ensuring that viewers would not suffer a loss of enjoyment when consuming media content by avoiding overlays obscuring items of importance. Consider claim 17, Zhang teaches a method comprising: receiving, by a computing device, an overlay object, comprising an object, for insertion into video content (Paragraph 0037 teaches videos that are streamed to a user can include additional content, e.g., provided by a content provider 108, that is overlaid on top of the original video stream); identifying, in one or more frames of a plurality of frames of the video content, and by using a machine-learned algorithm, one or more first content objects to remain at least partially unobscured by the overlay object (Paragraph 0012 teaches machine learning engines can identify important features within the video stream. Areas can be identified that encompass these important features, and then the overlaid content can be displayed outside of these identified areas. Paragraph 0013 teaches using machine learning to identify that fraction of the viewing area that contains the important content of the underlying video stream, and overlaying additional content outside of that fraction of the viewing area. Paragraph 0027 teaches identifying portions of video where overlaid content can be provided without occluding important video content. Automatically identify, within a set of frames of video, video features of video features types corresponding to important content. Paragraph 0028 teaches certain features of the video may be important in that overlaying content over those features would interfere with the viewing of the video. Paragraph 0039 teaches overlaying content on top of a video stream, while at the same time avoiding areas of the video screen that feature important content in the underlying video stream, e.g., areas in the original video stream that contain faces, text, or significant objects such as moving objects); identifying, based on the one or more first content objects, one or more regions of the one or more frames of the plurality of frames, where action of one or more second content objects is to occur (Paragraph 0003 teaches determining a location at which to position overlaid content during video display. Paragraph 0007 teaches determining location at which to position overlaid content during video display based on duration or number of frames over which the overlaid content is to be provided within the video. Paragraph 0028 teaches assigning different levels of importance to each detected feature type. For sports-related video a high importance may be assigned for not overlaying portions of the video where there is motion. Paragraph 0029 teaches based on the adjusted confidence values for the different feature types, the video system can identify one or more locations at which the overlay content can be displayed. Paragraph 0088 teaches determining based on the aggregated and adjusted confidence scores, a location at which to display the overlaid content during display of the video. Overlay location is determined based on desired display duration. Paragraph 0060 teaches other types of important features that a user may prefer to not be occluded can be moving objects. For example, objects that are in motion are generally more likely to convey important content to the viewer than static objects and are therefore generally less suitable to be occluded by overlaid content. The moving object detector 310e can detect moving objects by detecting movement between adjacent sampled frames, based on color-space differences between frames. Processing can be performed for each frame, other than first and last frames, using a previous, current, and next frame, with calculations producing a confidence score for the current frame. Paragraph 0067 teaches each feature type weight for a category can be, for example, a value between zero and one. A weight value of less than one, e.g., 0.5, for a feature type for a video category can indicate that the feature type is less important for the video category, a sports video category may have a weight of one for moving object features, a weight of one for text features, e.g., for scoreboards or statistics, and a weight of 0.8 for human face features. Paragraph 0073 teaches detected features, as represented by aggregated and adjusted confidence scores, may be situated at different locations within different frames that occupy a certain duration of the input video 302. Locations that include identified features can be locations that should not, if possible, be occluded by overlaid content, e.g., to ensure that important content of the underlying input video 302 is not occluded. As another example, an overlaid content location 320a can be a location which corresponds to confidence scores that are below a predetermined threshold. As locations with motion are taken into consideration when selecting a location for the overlay object, successive sequential frames are compared to identify motion, where regions with motion, can correspond to a second region that is identified. These regions are granted a higher importance/weight, which should not be occluded by overlaid content); determining, based on a first importance value corresponding to the one or more second content objects, a first time period during which the one or more second content objects should not be obstructed by one or more overlay objects (Paragraph 0028 teaches assigning different levels of importance to each detected feature type. For sports-related video a high importance may be assigned for not overlaying portions of the video where there is motion. Paragraph 0060 teaches other types of important features that a user may prefer to not be occluded can be moving objects. For example, objects that are in motion are generally more likely to convey important content to the viewer than static objects and are therefore generally less suitable to be occluded by overlaid content. The moving object detector 310e can detect moving objects by detecting movement between adjacent sampled frames, based on color-space differences between frames. Processing can be performed for each frame, other than first and last frames, using a previous, current, and next frame, with calculations producing a confidence score for the current frame. Paragraph 0067 teaches each feature type weight for a category can be, for example, a value between zero and one. A weight value of less than one, e.g., 0.5, for a feature type for a video category can indicate that the feature type is less important for the video category, a sports video category may have a weight of one for moving object features, a weight of one for text features, e.g., for scoreboards or statistics, and a weight of 0.8 for human face features. Paragraph 0073 teaches detected features, as represented by aggregated and adjusted confidence scores, may be situated at different locations within different frames that occupy a certain duration of the input video 302. Locations that include identified features can be locations that should not, if possible, be occluded by overlaid content, e.g., to ensure that important content of the underlying input video 302 is not occluded. For example, if the video category is sports, a weight value of 1 may be given for moving object features, where moving objects may not be occluded. Movement of moving objects contain motion between adjacent frames, in different locations that occupy a certain duration of the input video. Thus, during this duration of motion of the object, the object must not be obstructed by an overlay object); preventing output, in the video content and for the first time period, of an overlay object in the one or more regions; and after determining that the first time period has elapsed, causing, output of the video content with the overlay object inserted into the one or more regions (Paragraph 0003 teaches providing the overlaid content for display at the determined location in the video. Paragraph 0028, 0060 teaches for sports related video, a high importance may be assigned for not overlaying portions of the video where there is motion. Paragraph 0067 teaches sports category may have a weight of one for moving object features. Paragraph 0089 teaches video processing system 110 provides the overlaid content for display at the determined location in the video. Video processing system 110 can provide an overlaid content location 320a and a corresponding time offset 320 for an overlaid content item, and in some implementations, the overlaid content item itself, to the content platform 106. Paragraph 0090 teaches video processing system 110 and/or content platform 106 can display an overlaid content item on top of video content without occluding important video content. Paragraph 0071 teaches one or more existing content slots used to display overlay content on the input video 302. Paragraph 0073 teaches detected features, as represented by aggregated and adjusted confidence scores, may be situated at different locations within different frames that occupy a certain duration of the input video 302. An overlaid content location 320a can be determined so overlaid content that is displayed starting at a time offset 320b and for a specified duration of time that includes multiple frames is not positioned anywhere that the detected features are situated across the multiple frames. An identified overlaid content location 320a can be a location generally or substantially outside locations that include the identified features in a sequence of video frames corresponding to the input duration. Locations that include identified features can be locations that should not, if possible, be occluded by overlaid content, e.g., to ensure that important content of the underlying input video 302 is not occluded. As another example, an overlaid content location 320a can be a location which corresponds to confidence scores that are below a predetermined threshold. Paragraph 0078 teaches the overlay location identifier 318 can determine location(s) that correspond to low confidence scores that persist across a set of frames, for a location for display of overlay content for a set of frames equaling a desired duration. Paragraph 0080 teaches a content slot 506, which is located outside of colored areas corresponding to high confidence scores, corresponds to a recommended location for display of an overlaid content item for a specified duration starting at, or at least including, a time offset of the talk show video corresponding to the video frame 200. Based on importance, overlays may be prevented from being output over areas with important content, such as moving objects. Movement of moving objects contain motion between adjacent frames, in different locations that occupy a certain duration of the input video. Thus, during this duration of motion of the object, the object must not be obstructed by an overlay object. However, past the duration of the motion of the object, where the corresponding areas may no longer have movement, those areas become unimportant now as importance was determined based on motion of object(s). For areas having unimportant or less important content, areas where confidence score are below a threshold, outside of areas of high confidence scores, overlay content may be displayed. Thus, during times where objects have motion, overlay is prevented from display on those corresponding areas, but when there is no longer motion after a certain duration, those same areas become unimportant, and overlays may be displayed on those corresponding areas). Zhang does not explicitly teach where action is to occur, is where future action of one or more second content objects is predicted to occur; In an analogous art, Gupta teaches identifying, based on the one or more first content objects, one or more regions of the one or more frames, of the plurality of frames, where future action of one or more second content objects is predicted to occur (Abstract, Figs.1, Paragraph 0027-0028, 0034-0035, 0041, 0043; Fig.6, Paragraph 0107-0110). Therefore, it would have been obvious to a person of ordinary skill in the art to modify the system of Zhang to include identifying, based on the one or more first content objects, one or more regions of the one or more frames, of the plurality of frames, where future action of one or more second content objects is predicted to occur, as taught by Gupta, for the advantage of generating for display an overlay in a location that does not overlap with any of the first location, the second location, and the projected location (Gupta – Abstract, Paragraph 0110), better ensuring that viewers would not suffer a loss of enjoyment when consuming media content by avoiding overlays obscuring items of importance. Consider claims 2 and 19, Zhang and Gupta teach wherein the overlay object comprises one or more of: a uniform resource locator (URL), a hashtag, a quick response (QR) code, information about the video content, identification of an individual in the video content, identification of a sponsor of the video content, an animation, an advertisement, a chyron, a nameplate, a picture-in-picture, a timed text markup language (TTML) object, a timer, a ticker, a caption, or a graphic (Zhang - Paragraph 0026). Consider claims 3 and 20, Zhang and Gupta teach wherein the video content comprises one or more of: sports content (Zhang - Paragraph 0028), news content, streaming video, animated content, interviews, award ceremonies, or entertainment content. Consider claim 4, Zhang and Gupta teach wherein the one or more first content objects comprise one or more of: a player in a sporting event (Zhang - Paragraph 0044), an individual in news content, a character or actor in entertainment content, an individual in an interview or awards show, a score box, a chyron, items or regions of the video content referenced or used by individuals of the video content, or a portion of a frame showing an advertisement. Consider claim 5, Zhang and Gupta teach wherein the causing output of the video content with the overlay object inserted into the one or more regions comprises sending, to a second computing device, the video content with the overlay object inserted into the one or more regions (Zhang - Paragraph 0037; Paragraph 0073, 0078, 0080; Paragraph 0003, 0089, 0090). Consider claim 6, Zhang and Gupta teach wherein the preventing output of the overlay object in the one or more regions comprises causing display of the overlay object in one or more second regions of the one or more frames (Zhang - Paragraph 0008, 0011-0014). Consider claim 7, Zhang and Gupta teach further comprising: determining a plurality of content objects present in at least a portion of the plurality of frames (Zhang - Paragraph 0027; Fig.3, Paragraph 0051-0060); and selecting, from the plurality of content objects, the one or more first content objects (Zhang - Paragraph 0028-0029, 0067-0068). Consider claim 8, Zhang and Gupta teach further comprising: selecting, based on one or more different importance values for one or more different content objects, one or more second regions of the one or more frames, wherein the preventing output of the overlay object in the one or more regions comprises causing display of the overlay object in the one or more second regions (Zhang - Paragraph 0008, 0011-0014,0067-0076, 0078, 0080). Consider claim 10, Zhang and Gupta teach wherein the first importance value indicates a priority value for the one or more second content objects to remain unobscured (Zhang - Paragraph 0008, 0011-0014). Consider claim 12, Zhang and Gupta teach wherein the identifying the one or more first content objects associated with the highest importance value comprises: receiving, for the video content, one or more potential content objects and an importance value for each of the one or more potential content objects; determining one or more content objects, of the one or more potential content objects, in the video content; determining, based on the importance values of the one or more potential content objects, an importance value for each of the one or more content objects in the video content; and identifying the one or more first content objects by identifying, based on the determined importance values of the one or more content objects in the video content, the one or more content objects associated with the highest importance values (Zhang - Paragraph 0027-0029, 0067-0076, 0078, 0080). Consider claim 14, Zhang and Gupta teach wherein the identifying the one or more first content objects associated with the highest importance values comprises: determining a position for each of the one or more first content objects; and determining an importance value for each of the one or more first content objects (Zhang - Paragraph 0027-0029, 0067-0076, 0078, 0080). Consider claim 15, Zhang and Gupta teach further comprising: selecting the one or more regions based on a time period that the overlay object will at least partially obscure at least a portion of a plurality of frames of the video content (Zhang - Paragraph 0007, 0073-0075, 0078). Consider claim 16, Zhang and Gupta teach further comprising receiving one or more rule packages that comprise one or more of a time period and a shape of a region, of at least a portion of a plurality of frames of the video content, to obscure (Zhang - Paragraph 0073-0074, 0088). Claim(s) 9 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US 2022/0368979), in view of Gupta et al. (US 2019/0349626), and further in view of Pandhare et al. (US 2023/0149819). Consider claim 9, Zhang and Gupta teach further comprising: identifying the one or more first content objects by: identifying using a machine-learned algorithm, a content object that appears, at least partially and for a period of time, in the plurality of frames of the video content (Zhang - Paragraph 0012-0013, 0038-0039, 0055), but do not explicitly teach a machine-learned algorithm trained using one or more of logistic regression, multi nominal logistics regression, linear regression, support vector machines, naive Bayes, decision trains, k nearest neighbors, random forest, boosting, k-means, or hierarchical clustering. In an analogous art, Pandhare teaches a machine-learned algorithm trained using one or more of logistic regression, multi nominal logistics regression, linear regression (Paragraph 0153), support vector machines, naive Bayes, decision trains, k nearest neighbors, random forest, boosting, k-means, or hierarchical clustering. Therefore, it would have been obvious to a person of ordinary skill in the art to modify the system of Zhang and Gupta to include a machine-learned algorithm trained using one or more of logistic regression, multi nominal logistics regression, linear regression, support vector machines, naive Bayes, decision trains, k nearest neighbors, random forest, boosting, k-means, or hierarchical clustering, as taught by Pandhere, for the advantage of utilizing algorithm(s) that is relatively straightforward to understand and implement, is computationally efficient, versatile, and scalable. Consider claim 18, Zhang and Gupta teach the machine-learned algorithm, to identify content objects to remain at least partially unobscured by overlay objects (Zhang - Paragraph 0012-0013, 0038-0039, 0055), but do not explicitly teach wherein the machine-learned algorithm was trained, using one or more of logistic regression, multi nominal logistics regression, linear regression, support vector machines, naive Bayes, decision trains, k nearest neighbors, random forest, boosting, k-means, or hierarchical clustering. In an analogous art, Pandhare teaches wherein a machine-learned algorithm was trained, using one or more of logistic regression, multi nominal logistics regression, linear regression (Paragraph 0153), support vector machines, naive Bayes, decision trains, k nearest neighbors, random forest, boosting, k-means, or hierarchical clustering. Therefore, it would have been obvious to a person of ordinary skill in the art to modify the system of Zhang and Gupta to include wherein a machine-learned algorithm was trained, using one or more of logistic regression, multi nominal logistics regression, linear regression, support vector machines, naive Bayes, decision trains, k nearest neighbors, random forest, boosting, k-means, or hierarchical clustering, as taught by Pandhere, for the advantage of utilizing algorithm(s) that is relatively straightforward to understand and implement, is computationally efficient, versatile, and scalable. Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US 2022/0368979), in view of Gupta et al. (US 2019/0349626), and further in view of Walker et al. (US 8,281,339). Consider claim 13, Zhang and Gupta teach further comprising: determining, based on a size of the overlay object, a size of a region to be obscured by the overlay object; and selecting, based on an insertion period and the size of the region to be obscured, the one or more regions (Paragraph 0007, 0074-0075; Paragraph 0027-0029, 0067-0076), but does not explicitly teach determining, based on shape of the overlay object, shape of a region, and the shape of the region. In an analogous art, Walker teaches determining, based on shape of the overlay object, shape of a region, and the shape of the region (Col 25: lines 20-50). Therefore, it would have been obvious to a person of ordinary skill in the art to modify the system of Zhang and Gupta to include determining, based on shape of the overlay object, shape of a region, and the shape of the region, as taught by Walker, for the advantage of preventing content from being obscured (Walker – Col 25: lines 49-50), providing finer granularity and control in the processing, selection, and display of content. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JASON K LIN whose telephone number is (571)270-1446. The examiner can normally be reached on Monday-Friday 9AM-5PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Brian Pendleton can be reached on 571-272-7527. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JASON K LIN/Primary Examiner, Art Unit 2425
Read full office action

Prosecution Timeline

Show 7 earlier events
Sep 23, 2025
Interview Requested
Oct 02, 2025
Examiner Interview Summary
Oct 02, 2025
Applicant Interview (Telephonic)
Oct 03, 2025
Request for Continued Examination
Oct 07, 2025
Response after Non-Final Action
Jan 09, 2026
Non-Final Rejection mailed — §103
Feb 04, 2026
Response Filed
Jun 24, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12701274
CONTENT DISTRIBUTION AND OPTIMIZATION SYSTEM AND METHOD FOR DERIVING NEW METRICS AND MULTIPLE USE CASES OF DATA CONSUMERS USING BASE EVENT METRICS
4y 7m to grant Granted Aug 04, 2026
Patent 12695938
BROADCAST RECEIVING APPARATUS, BROADCAST RECEIVING METHOD, AND CONTENTS OUTPUTTING METHOD
1y 5m to grant Granted Jul 28, 2026
Patent 12689798
METHODS AND APPARATUS TO DETERMINE A NUMBER OF PEOPLE IN AN AREA
2y 4m to grant Granted Jul 21, 2026
Patent 12677042
METHOD, APPARATUS, ELECTRONIC DEVICE, AND STORAGE MEDIUM FOR PROCESSING LIVE STREAMING INFORMATION
3y 10m to grant Granted Jul 07, 2026
Patent 12621504
TECHNIQUES FOR CACHING MEDIA CONTENT WHEN STREAMING LIVE EVENTS
3y 1m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
49%
Grant Probability
83%
With Interview (+33.9%)
3y 8m (~1y 8m remaining)
Median Time to Grant
High
PTA Risk
Based on 460 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month