Prosecution Insights
Last updated: August 06, 2026
Application No. 18/702,105

VIDEO IMPLANTATION METHOD, APPARATUS, DEVICE AND COMPUTER-READABLE STORAGE MEDIUM

Final Rejection §103
Filed
Apr 17, 2024
Priority
Oct 21, 2021 — CN 202111227816.4 +1 more
Examiner
LIN, JASON K
Art Unit
2425
Tech Center
2400 — Computer Networks
Assignee
Xingheshixiao (Beijing) Technology Co. Ltd.
OA Round
4 (Final)
49%
Grant Probability
Moderate
5-6
OA Rounds
1y 5m
Est. Remaining
83%
With Interview

Examiner Intelligence

Grants 49% of resolved cases
49%
Career Allowance Rate
225 granted / 459 resolved
-9.0% vs TC avg
Strong +34% interview lift
Without
With
+33.7%
Interview Lift
resolved cases with interview
Typical timeline
3y 8m
Avg Prosecution
21 currently pending
Career history
488
Total Applications
across all art units

Statute-Specific Performance

§101
5.6%
-34.4% vs TC avg
§103
63.4%
+23.4% vs TC avg
§102
14.5%
-25.5% vs TC avg
§112
8.8%
-31.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 459 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION This office action is responsive to application No. 18/702,105 filed on 12/18/2025. Claim(s) 2, 4, and 8-10 have been cancelled. Claim(s) 1, 3, 5-7, and 11-12 is/are pending and have been examined. Claim Objections Claim(s) 1 is/are objected to because of the following informalities: Claim 1 recites: “performing semantic analysis and/or content analysis on the low code source video through a video port provided by a publisher, it being unnecessary to obtain the complete source low code video when the low code source video being analyzed” Please amend to: --performing semantic analysis and/or content analysis on the low code source video through a video port provided by a publisher, it being unnecessary to obtain the complete source video when the low code source video being analyzed-- Response to Arguments Applicant's arguments filed 05/25/2026 have been fully considered but they are not persuasive. A) Applicants assert on P.5 of 8 that “1. Core Technical Difference. A high resolution version of those identified parts of the video is then sent to the specialist for visual object insertion. The specialist may then return the modified parts of the video and the content owner create a final version of the high-resolution video by replacing the relevant parts of the high resolution video with the modified parts.” The Examiner appreciates the core technical difference assert by Applicant. However, the Examiner respectfully disagrees. Although Applicant makes a distinction on how Applicant’s invention may be technically different than Fauqueur or Srinivasan, the assert core technical difference “A high resolution version of those identified parts of the video is then sent to the specialist for visual object insertion. The specialist may then return the modified parts of the video and the content owner create a final version of the high-resolution video by replacing the relevant parts of the high resolution video with the modified parts” is not claimed anywhere in the claims, therefore, claims need not be bounded by the explanation of the technical difference made by the Applicant. If Applicant intends claims to be bounded by the core technical difference, the assertions made by Applicant should be incorporated as limitation(s) into the claimed invention. B) Applicant assert on P.5 of 8 that “…neither Fauqueur nor Srinivasan discloses or suggests such as generating object description information according to the visual object and the source video clip corresponding to the one or more frames+ sending the visual object and the object description information to the publisher, so that the publisher overlays the visual object on the frame corresponding to the one or more frames in the source video data in a mask manner according to the object description information; or sending the masked visual object and the object description information to the publisher, so that the publisher implants the masked visual object into the frame corresponding to one or more frames in a rendering fusion manner according to the object description information to obtain the final video.” In response, the Examiner respectfully disagrees. Please see rejection of these claim limitation(s) by the combination of Srinivasan and Fauqueur in the Office Action below. C) Applicant asserts on P.6 of 8 that “Fauqueur similarly fails to disclose generating object description information according to the visual object and the source video clip corresponding to the one or more frames+ sending the visual object and the object description information to the publisher, so that the publisher overlays the visual object on the frame corresponding to the one or more frames in the source video data in a mask manner according to the object description information; or sending the masked visual object and the object description information to the publisher, so that the publisher implants the masked visual object into the frame corresponding to one or more frames in a rendering fusion manner according to the object description information to obtain the final video.” In response, the Examiner respectfully disagrees. These claim limitation(s) are taught by Fauqueur. Please see Office Action below. D) Applicant asserts on P.6 of 8 that “The claimed technical solution achieves advantages neither taught nor suggested by the prior art, particularly by requiring only the sending of the visual object and the object description information to the publisher to obtain the final video.” In response, the Examiner respectfully disagrees. Although claim recites “sending the visual object and the object description information to the publisher”, the absence of the limitation of sending video to the publisher does not preclude the sending of video to the publisher. If Applicant intends it as such, Applicant must include in the claim only sending the visual object and the object description information to the publisher, and not sending the video to the publisher. Based on the above and the Office Action below. Although Applicant does bring up some technical distinctions between Applicant’s invention and the cited references of record. Applicant has not fully claimed these technical distinctions in the claim(s), therefore, Srinivasan and Fauqueur continue to teach the claimed limitation(s). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1, 5-7, and 11-12 is/are rejected under 35 U.S.C. 103 as being unpatentable over in view of Srinivasan et al. (US 2020/0372937) in view of Fauqueur et al. (US 2014/0147096). Consider claim 1, Srinivasan teaches a video implantation method, comprising steps of: analyzing a low code source video and recognizing one or more frames in which a visual object can be implanted (Paragraph 0038 teaches a source hub 102 comprising a video data analysis module, which performs pre-analysis in relation to source video data. Paragraph 0039 teaches pre-analysis may be fully automated in that it does not involve any human intervention. Paragraph 0040-0050 teaches the different types of pre-analysis that may be performed. Paragraph 0051 teaches source hub analyses the source video data to find regions within the source video material which are suitable for insertion of one or more additional visual objects in the image content of the source video. Paragraph 0071 teaches in step 320, the source hub 302 analyses the low-resolution video material to identify one or more “insertion zones” that correspond to one or more regions within the image contents of the video material that are suitable for insertion one or more additional visual objects. Paragraph 0072 teaches remainder of low-resolution video material is analysed to identify if the insertion zone appears in one or more other shots in the low-resolution video material); acquiring a source video clip corresponding to the one or more frames (Paragraph 0074 teaches source hub 102 obtains high-resolution video data comprises a second plurality of frames of the video material that comprise at least the selected frames of the video material. The second plurality of frames are a smaller sub-set of the first plurality of frames. such that the source hub 102 obtains only a part of the high-resolution video material. The source hub 102 does not obtain the entire high-resolution video material. In some examples, the second plurality of frames may consist of only the selected frames of the video material. In other examples, the second plurality of frames may consist of the selected frames and also some additional frames of the video material); the one or more frames being obtained by the following steps: analyzing the low code source video and recognizing one or more frames in which the visual object can be implanted (Paragraph 0038 teaches a source hub 102 comprising a video data analysis module, which performs pre-analysis in relation to source video data. Paragraph 0039 teaches pre-analysis may be fully automated in that it does not involve any human intervention. Paragraph 0040-0050 teaches the different types of pre-analysis that may be performed. Paragraph 0051 teaches source hub analyses the source video data to find regions within the source video material which are suitable for insertion of one or more additional visual objects in the image content of the source video. Paragraph 0071 teaches in step 320, the source hub 302 analyses the low-resolution video material to identify one or more “insertion zones” that correspond to one or more regions within the image contents of the video material that are suitable for insertion one or more additional visual objects. Paragraph 0072 teaches remainder of low-resolution video material is analysed to identify if the insertion zone appears in one or more other shots in the low-resolution video material); the analyzing the low code source video (Paragraph 0040-0051, 0071-0072) comprises: performing semantic analysis and/or content analysis on the low code source video through a video port provided by a publisher, it being unnecessary to obtain the complete source low code video when the low code source video being analyzed; low code source video, low code source vide; determine one or more frames in which the visual object can be implanted (Paragraph 0035 teaches source hub may retrieve source video data as one or more digital files, supplied, for example, over a high-seed computer network, via the network, etc. Paragraph 0036 teaches source video data may be provided by a distributor or content owner. In the case of live video, source video would be provided on an on-going basis, so the complete source video would not be provided as it is not yet completed, thus analysis of the source video would occur on available segments/chunks of live video that is provided from the distributor. Paragraph 0038-0051 teaches source hub 102 comprising a video data analysis module, that may perform may different types of pre-analysis on the source video data, to find regions within the source video material which are suitable for insertion of one or more additional visual objects in the image content of the source video. Paragraph 0062 teaches video data of the distributor, where type of content may be live, VoD, etc. Paragraph 0068 teaches source hub 102 obtains source video data from a source video data holding entity 304, for example, the distributor or content producer. Paragraph 0071 teaches in step 320, the source hub 302 analyses the low-resolution video material to identify one or more “insertion zones” that correspond to one or more regions within the image contents of the video material that are suitable for insertion one or more additional visual objects); Srinivasan does not explicitly teach generating object description information according to the visual object and the source video clip corresponding to the one or more frames, the object description information being used to describe the implantation position of the visual object in the one or more frames and the specific information of the one or more frames, after the semantic analysis and/or content analysis is performed on the source video, determining one or more videos in the source video that satisfy a preset requirement, the preset requirement being associated with the visual object; and analyzing the one or more videos to determine one or more frames in which the visual object can be implanted, wherein the method further comprises: sending the visual object and the object description information to the publisher, so that the publisher overlays the visual object on the frame corresponding to the one or more frames in the source video data in a mask manner according to the object description information; or sending the masked visual object and the object description information to the publisher, so that the publisher implants the masked visual object into the frame corresponding to one or more frames in a rendering fusion manner according to the object description information to obtain the final video. In an analogous art, Fauqueur teaches generating object description information according to the visual object and the source video clip corresponding to the one or more frames, the object description information being used to describe the implantation position of the visual object in the one or more frames and the specific information of the one or more frames (Paragraph 0050 teaches frames in the video data. Paragraph 0101 teaches an embed project for adding one or more additional video objects to one or more segments identified in step 2f. Paragraph 0123-0124 teaches frames covered by the embed project and timecodes corresponding to frames in the embed sequence. Paragraph 0162 teaches video file data sent may be a package including (1) a project file defining the tracking, masking, appearance modeling, embed artwork, and other data (such as effects data); (2) the embed artwork; and (3) some or all of the embed project metadata. Paragraph 0143 teaches tracking involves tracking the position of the virtual product, as it will appear in the embedded sequence. Tracking is used to determine the horizontal and vertical position of the virtual product on each frame of the embed sequence in which the product is to be placed), after semantic analysis and/or content analysis is performed on source video, determining one or more videos in the source video that satisfy a preset requirement, the preset requirement being associated with a visual object; and analyzing the one or more videos to determine one or more frames in which the visual object can be implanted (Paragraph 0102-0104), wherein the method further comprises: sending the visual object and the object description information to the publisher, so that the publisher overlays the visual object on the frame corresponding to the one or more frames in the source video data in a mask manner according to the object description information (Paragraph 0162, 0185-0192); or sending the masked visual object and the object description information to the publisher, so that the publisher implants the masked visual object into the frame corresponding to one or more frames in a rendering fusion manner according to the object description information to obtain the final video. Therefore, it would have been obvious to a person of ordinary skill in the art to modify the system of Srinivasan to include generating object description information according to the visual object and the source video clip corresponding to the one or more frames, the object description information being used to describe the implantation position of the visual object in the one or more frames and the specific information of the one or more frames, after semantic analysis and/or content analysis is performed on source video, determining one or more videos in the source video that satisfy a preset requirement, the preset requirement being associated with a visual object; and analyzing the one or more videos to determine one or more frames in which the visual object can be implanted, wherein the method further comprises: sending the visual object and the object description information to the publisher, so that the publisher overlays the visual object on the frame corresponding to the one or more frames in the source video data in a mask manner according to the object description information; or sending the masked visual object and the object description information to the publisher, so that the publisher implants the masked visual object into the frame corresponding to one or more frames in a rendering fusion manner according to the object description information to obtain the final video, as taught by Fauqueur, for the advantage of better identifying suitable segments, as not all of the identified segments of the source video data are, in fact, suitable for product placement, thus, not all of the identified segments are selected for digital product placement (Fauqueur – Paragraph 0102), allowing the system to best select segment(s) that may better suited contextually, and of greater interest (Fauqueur – Paragraph 0103), allowing the system to be easily convey and direct information detailing produce placement into content. Consider claim 5, Srinivasan and Fauqueur teach wherein, the analyzing the one or more videos to determine one or more frames in which the visual object can be implanted comprises: analyzing the one or more videos to determine a region of interest suitable for implantation of the visual object (Fauqueur - Paragraph 0053-0058; Srinivasan – Paragraph 0040-0050, 0071); and determining the frame where the region of interest is located as the one or more frames (Fauqueur - Paragraph 0099-0104, 0123). Consider claim 6, Srinivasan and Fauqueur teach wherein, the acquiring a source video clip corresponding to the one or more frames comprises: acquiring the frame corresponding to the one or more frames in high code rate source video data (Srinivasan - Paragraph 0074 teaches source hub 102 obtains high-resolution video data comprises a second plurality of frames of the video material that comprise at least the selected frames of the video material. The second plurality of frames are a smaller sub-set of the first plurality of frames. such that the source hub 102 obtains only a part of the high-resolution video material. The source hub 102 does not obtain the entire high-resolution video material. In some examples, the second plurality of frames may consist of only the selected frames of the video material. In other examples, the second plurality of frames may consist of the selected frames and also some additional frames of the video material); and the implanting the visual object into the source video clip corresponding to the one or more frames and generating one or more output videos comprises: implanting the visual object into the frame corresponding to the one or more frames in the high code rate source video data, and generating one or more output videos (Srinivasan - Paragraph 0075 teaches in step 350, the one or more additional visual objects are embedded into the selected frames of the high-resolution source video data. Paragraph 0076 teaches in step 360, output video data is created. The output video data comprises the selected frames of the high-resolution video material with the embedded one or more additional visual objects). Consider claim 7, Srinivasan and Fauqueur teach wherein, the acquiring the frame corresponding to the one or more frames in high code rate source video data further comprises: acquiring the frame corresponding to the one or more frames in the high code rate source video data according to a preset security frame strategy; wherein the preset security frame strategy is used to indicate the respective number of supplementary frames of the one or more frames (Srinivasan – Paragraph 0074, 0085-0086; Paragraph 0080-0082). Consider claim 11, Srinivasan and Fauqueur teach an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein, the memory has instructions stored thereon that can be executed by the at least one processor, and the instructions enable, when executed by the at least processor, the at least one processor to execute the method according to claim 1 (Srinivasan - Paragraph 0101-0105). Consider claim 12, Srinivasan and Fauqueur teach a non-transient computer-readable storage medium having computer instructions stored thereon that are configured to cause a computer to execute the method according to claim 1 (Srinivasan - Paragraph 0101-0105). Claim(s) 3 is/are rejected under 35 U.S.C. 103 as being unpatentable over in view of Srinivasan et al. (US 2020/0372937), in view of Fauqueur et al. (US 2014/0147096), and further in view of Loheide et al. (US 2021/0211750). Consider claim 3, Srinivasan and Fauqueur teach wherein, the generating object description information according to the visual object and the source video clip corresponding to the one or more frames comprises: analyzing a region of interest suitable for implantation of the visual object in the source video clip corresponding to the one or more frames to determine the object description information (Fauqueur – Paragraph 0054, 0058, 0143, 0162); and storing source videos by the publisher (Srinivasan – Paragraph 0062-0063); and the analyzing a source video and recognizing one or more frames in which a visual object can be implanted comprises: analyzing source video and recognizing one or more frames in which the visual object can be implanted (Fauqueur - Paragraph 0046, 0050, 0101, 0123-0124; Srinivasan – 0063-0064). Srinivasan and Fauqueur do not explicitly teach storing multiple versions of source videos, each version of source video being different in code rate and/or language version; and the analyzing a source video comprises: analyzing any version of source video among the multiple versions of source videos. In an analogous art, Loheide teaches storing multiple versions of source videos, each version of source video being different in code rate and/or language version (Paragraph 0088 teaches plurality of media segments may be processed and stored at various quality levels, and content encryption modes. For example, the media segment may be stored in a plurality of quality levels, for example, high definition (HD), high dynamic range (HDR) video, or different quality levels in accordance with specified pixel resolutions, bitrates, resolutions, bandwidths, frame rates, and/or sample frequencies); and the analyzing a source video comprises: analyzing any version of source video among the multiple versions of source videos (Paragraph 0088 teaches media segments stored at plurality of different quality levels. Pre-encoded media assets may be re-used to create new channels, program streams, without requiring to re-encode a selected media asset. Paragraph 0087 teaches insertion of live content, pre-stored media content, pre-encoded media assets, and/or the like, may be driven by real time or near-real time content context analysis. As any version of video may be retrieved to create channel and/or stream, content analysis, may be performed on any version of video, in order to provide for appropriate insertion of content). Therefore, it would have been obvious to a person of ordinary skill in the art to modify the system of Srinivasan and Fauqueur to include storing multiple versions of source videos, each version of source video being different in code rate and/or language version; and the analyzing a source video comprises: analyzing any version of source video among the multiple versions of source videos, as taught by Loheide, for the advantage of providing the network provider with the capability to not only provide new channel offerings in cost-effective manner, but also provide enhanced viewer experience to increase their appeal in order to gain a wider audience (Loheide – Paragraph 0013), providing greater flexibility and adaptability in processing and serving various different client devices, quickly and effectively. Conclusion THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JASON K LIN whose telephone number is (571)270-1446. The examiner can normally be reached on Monday-Friday 9AM-5PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Brian Pendleton can be reached on 571-272-7527. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see https://ppair-my.uspto.gov/pair/PrivatePair. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JASON K LIN/Primary Examiner, Art Unit 2425
Read full office action

Prosecution Timeline

Show 1 earlier event
Mar 03, 2025
Non-Final Rejection mailed — §103
Jun 02, 2025
Response Filed
Jun 18, 2025
Final Rejection mailed — §103
Dec 18, 2025
Request for Continued Examination
Jan 08, 2026
Response after Non-Final Action
Feb 24, 2026
Non-Final Rejection mailed — §103
May 25, 2026
Response Filed
Jun 18, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12701274
CONTENT DISTRIBUTION AND OPTIMIZATION SYSTEM AND METHOD FOR DERIVING NEW METRICS AND MULTIPLE USE CASES OF DATA CONSUMERS USING BASE EVENT METRICS
4y 7m to grant Granted Aug 04, 2026
Patent 12695938
BROADCAST RECEIVING APPARATUS, BROADCAST RECEIVING METHOD, AND CONTENTS OUTPUTTING METHOD
1y 5m to grant Granted Jul 28, 2026
Patent 12689798
METHODS AND APPARATUS TO DETERMINE A NUMBER OF PEOPLE IN AN AREA
2y 4m to grant Granted Jul 21, 2026
Patent 12677042
METHOD, APPARATUS, ELECTRONIC DEVICE, AND STORAGE MEDIUM FOR PROCESSING LIVE STREAMING INFORMATION
3y 10m to grant Granted Jul 07, 2026
Patent 12621504
TECHNIQUES FOR CACHING MEDIA CONTENT WHEN STREAMING LIVE EVENTS
3y 1m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
49%
Grant Probability
83%
With Interview (+33.7%)
3y 8m (~1y 5m remaining)
Median Time to Grant
High
PTA Risk
Based on 459 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month