Prosecution Insights
Last updated: October 01, 2026
Application No. 18/915,204

MEDIA CONTENT BOUNDARY-AWARE ENCODING

Non-Final OA §103
Filed
Oct 14, 2024
Priority
Jun 30, 2022 — continuation of 12/149,709
Examiner
CASTRO, ALFONSO
Art Unit
2421
Tech Center
2400 — Computer Networks
Assignee
Amazon Technologies Inc.
OA Round
3 (Non-Final)
51%
Grant Probability
Moderate
3-4
OA Rounds
1y 8m
Est. Remaining
70%
With Interview

Examiner Intelligence

Grants 51% of resolved cases
51%
Career Allowance Rate
230 granted / 451 resolved
-7.0% vs TC avg
Strong +19% interview lift
Without
With
+19.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 8m
Avg Prosecution
28 currently pending
Career history
490
Total Applications
across all art units

Statute-Specific Performance

§101
5.8%
-34.2% vs TC avg
§103
72.1%
+32.1% vs TC avg
§102
4.5%
-35.5% vs TC avg
§112
10.1%
-29.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 451 resolved cases

Office Action

§103
DETAILED ACTION Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 8/11/2026 has been entered. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s arguments, see Remarks pg. 9, filed 7/27/2026, with respect to the status of the claims and the interview summary are hereby acknowledged. Applicant’s arguments, see Remarks pg. 9, filed 7/27/2026, with respect to the non-statutory double patenting rejection has been fully considered and is hereby acknowledged. The filing of the terminal disclaimer is acknowledged and said rejection is withdrawn. The applicant’s Remarks, see pg. 9, filed 7/27/2026, with respect to the rejection(s) of claim(s) 1-20 under 35 U.S.C. 103 have been fully considered. The examiner acknowledges that applicant’s arguments are directed to newly amended limitations not previously presented. Therefore, the examiner will set forth a new grounds of rejection in order to address the new limitations. With respect to applicant’s arguments, the applicant argues that “[t]he cited references do not disclose a "CV/ML temporal boundary report log that is based at least in part on the media content.”” In particular, applicant argues the following: Amended claim 1 recites, in part, "utilizing a boundary generation process on the media content to determine a computer vision/machine learning (CV /ML) temporal boundary report that includes at least one of the first set of temporal boundaries or the second set of temporal boundaries and a CV/ML temporal boundary report log that is based at least in part on the media content." On page 17, the Office Action concedes that Mao "teaches all the elements of the claim except a report as claimed" and relies on Zhang to supply the recitations absent from Mao. The Office Action states that "Zhang teaches generating a set of candidate breakpoints in a media item and using a machine learning model to score the candidate breakpoints and iteratively select subsets of the candidate breakpoints to be a final set of breakpoints," and that "Zhang discloses a target set of boundaries in the form of a final set of breakpoints element 240 and 0-240 in figure 2." Zhang at paragraph [0011], on which the Office Action relies, discloses techniques that "programmatically evaluate candidate locations within media items at which breakpoint can be inserted, and generate a list of breakpoints at the locations within the media items that are predicted to be less disruptive to playback of the media items." Zhang at paragraph [0022] further discloses that "the media item includes (or references) a list of breakpoints that have been generated for the media item," the breakpoints specifying "timestamps within the duration of the media item where playback of the media item can be halted, and where digital components can be presented." Zhang at paragraph [0029] discloses a response that includes data for presenting a media item "along with a list of breakpoints (BP_ I-BP_ 2) that have been defined for that corresponding media item." Element 240 of Figure 2 of Zhang is described at paragraph [0053] of Zhang as determining "[a] final set of breakpoints from among the breakpoints in the subset of candidate breakpoints." To clarify, Zhang does not disclose "a CV/ML temporal boundary report log that is based at least in part on the media content," as recited in claim 1, and does not disclose a CV/ML temporal boundary report that includes both at least one of the recited sets of temporal boundaries and the recited CV/ML temporal boundary report log. On page 17, the Office Action identifies the final set of breakpoints of Zhang as corresponding to a "target set of boundaries." However, the Office Action identifies no disclosure of Zhang, and no disclosure of Mao, alleged to correspond to the recited "CV/ML temporal boundary report log." The examiner respectfully disagrees. In response to applicant’s argument, the test for obviousness is not whether the features of a secondary reference may be bodily incorporated into the structure of the primary reference; nor is it that the claimed invention must be expressly suggested in any one or all of the references. Rather, the test is what the combined teachings of the references would have suggested to those of ordinary skill in the art. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981). Additionally, on the issue of obviousness, the Supreme Court stated the analysis of a rejection on obviousness grounds need not seek out precise teachings directed to the specific subject matter of the challenged claim, for a court can take account of the inferences and creative steps that a person of ordinary skill in the art would employ. See KSR International Co. v. Teleflex Inc., 550 U.S. 398, 418, 82 USPQ2d 1385 (2007). The obvious analysis cannot be confined by a formalistic conception of the words teaching, suggestion, and motivation. Id. at 419. Further, the Court stated that common sense teaches, however, that familiar items may have obvious uses beyond their primary purposes, and in many cases a person of ordinary skill will be able to fit the teachings of multiple patents together like pieces of a puzzle. Id. at 420. Based on the principles of law as discussed above, Zhang paragraph [0023] teaches the following: As described in more detail below, the breakpoints for media items can be selected in a manner that reduces the disruption to playback of the media items. More specifically, a subset of all potential candidate breakpoints for a given media item can be selected based on their level of disruptiveness. The level of disruptiveness for each candidate breakpoint can be assessed, for example, based on an output from a machine learning model that has been trained to predict the disruptiveness of breakpoints based on characteristics of the given media item at a time of the candidate breakpoint and based on characteristics of frames of the given media item that are within a specified distance (e.g., amount of time or number of frames) of the time of the candidate breakpoint. The subset of the candidate breakpoints having the lowest predicted level of disruptiveness (e.g., according to the scores output by the machine learning model) can then be ranked based on one or more criteria, such as the relative proximity of each candidate breakpoint to other candidate breakpoints, in combination with the predicted level of disruptiveness. This ranking can then be used to select a threshold number of highest ranked breakpoints that will be used as final breakpoints for the given media item. Applicant has amended the limitations of the independent claims to recite the representative amendment of “merging the first set of temporal boundaries and the second set of temporal boundaries to generate a target set of temporal boundaries, the target set of temporal boundaries including one or more selected temporal boundaries selected from among the first set of temporal boundaries and the second set of temporal boundaries, wherein the one or more selected temporal boundaries are associated with higher accuracy levels than one or more unselected temporal boundaries from among the first set of temporal boundaries and the second set of temporal boundaries.” Emphasis Ours. Zhang renders obvious the claimed “higher accuracy levels” wherein, as discussed above, Zhang teaches utilizing machine learning, comprising a machine learning model can include a bi-directional gradient recurring unit and fully connected neural network layers, to generate a target set of temporal boundaries wherein the selected target set of temporal boundaries were selected based on scores and rankings that relate to being (see Zhang para 23, 32 - breakpoint management system 160 implements a combination of machine learning models and search techniques to process and analyze each media item to determine a set of breakpoints that limit the disruption to playback of the media item). Zhang teaches determining the optimal time to present particular breakpoints in the media based on scores and ranking and which corresponds to “higher accuracy levels than one or more unselected temporal boundaries from among the first set of temporal boundaries and the second set of temporal boundaries” because the chose breakpoints are a more accurate position to optimally present a breakpoint for presentation of content. All things considered, the applicant’s arguments regarding the teachings of Zhang are not persuasive. Therefore, a new grounds of rejection is being set forth in order to take into consideration the newly amended limitations. Claim Rejections - 35 USC § 103 The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action. Claim(s) 1-3, 8, 11-12, 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Mao; Weidong et al. US 20200288149 A1 (hereafter Mao) and in further view of Zhang; Wenbo et al. US 20210390130 A1 (hereafter Zhang). Regarding clam 1, “a method comprising: determining media content that includes at least one of video content or audio content; determining a first set of temporal boundaries indicating one or more first locations associated with a first portion of the media content; determining a second set of temporal boundaries indicating one or more second locations associated with a second portion of the media content; merging the first set of temporal boundaries and the second set of temporal boundaries to generate a target set of temporal boundaries, the target set of temporal boundaries including one or more selected temporal boundaries selected from among the first set of temporal boundaries and the second set of temporal boundaries, wherein the one or more selected temporal boundaries are associated with higher accuracy levels than one or more unselected temporal boundaries from among the first set of temporal boundaries and the second set of temporal boundaries; utilizing a boundary generation process on the media content to determine a computer vision/machine learning (CV/ML) temporal boundary report that includes at least one of the first set of temporal boundaries or the second set of boundaries and a CV/ML temporal boundary report log that is based at least in part on the media content; and encoding, based at least in part on the target set of temporal boundaries and the CV/ML temporal boundary report, the media content as encoded media content” Mao teaches para [0037-0045] identifying boundaries of using metadata included with the video wherein number of encoders available to encode scenes and which encoding parameters may be used by specific encoders may be determined; Target resolutions and/or bit rates for subsequent transmission of scenes may be determined; the computing device may determine that each scene should be encoded; See para 33-36 regarding merging boundaries. Mao ([0044] discloses identifying visual elements using a machine learning algorithm). Additionally, with respect to the modification of the term boundary to “temporal boundary,” the examiner notes that the broadest reasonable interpretation of the term “temporal” includes the dictionary definition “of or relating to time as opposed to eternity” and “of or relating to time as distinguished from space; of or relating to the sequence of time or to a particular time.” In view of the broadest reasonable interpretation of the term temporal, the prior art of record to Mao and Zhang render the limitation obvious. The prior art identifies boundaries and/or breakpoints with respect to a timeline of the presentation of video content frames as opposed to spatial content displayed in the video scenes. Mao does not disclose a report as claimed and also does not reference “higher accuracy levels” with respect to “the target set of temporal boundaries including one or more selected temporal boundaries selected from among the first set of temporal boundaries and the second set of temporal boundaries, wherein the one or more selected temporal boundaries are associated with higher accuracy levels than one or more unselected temporal boundaries from among the first set of temporal boundaries and the second set of temporal boundaries.” Regarding the deficiency of Mao, Zhang teaches generating a set of candidate breakpoints in a media item, and using a machine learning model to score the candidate breakpoints and iteratively select subsets of the candidate breakpoints to be a final set of breakpoints, which are stored as part of a bitstream. See para 11, 22, 29 regarding a list and figure 2. Zhang discloses a target set of boundaries in the form of a final set of breakpoints element 240 and 0-240 in figure 2. With respect to “higher accuracy levels” with respect to “the target set of temporal boundaries including one or more selected temporal boundaries selected from among the first set of temporal boundaries and the second set of temporal boundaries, wherein the one or more selected temporal boundaries are associated with higher accuracy levels than one or more unselected temporal boundaries from among the first set of temporal boundaries and the second set of temporal boundaries.” Zhang paragraph [0023] teaches the following: As described in more detail below, the breakpoints for media items can be selected in a manner that reduces the disruption to playback of the media items. More specifically, a subset of all potential candidate breakpoints for a given media item can be selected based on their level of disruptiveness. The level of disruptiveness for each candidate breakpoint can be assessed, for example, based on an output from a machine learning model that has been trained to predict the disruptiveness of breakpoints based on characteristics of the given media item at a time of the candidate breakpoint and based on characteristics of frames of the given media item that are within a specified distance (e.g., amount of time or number of frames) of the time of the candidate breakpoint. The subset of the candidate breakpoints having the lowest predicted level of disruptiveness (e.g., according to the scores output by the machine learning model) can then be ranked based on one or more criteria, such as the relative proximity of each candidate breakpoint to other candidate breakpoints, in combination with the predicted level of disruptiveness. This ranking can then be used to select a threshold number of highest ranked breakpoints that will be used as final breakpoints for the given media item. Applicant’s claim recites “the target set of temporal boundaries including one or more selected temporal boundaries selected from among the first set of temporal boundaries and the second set of temporal boundaries, wherein the one or more selected temporal boundaries are associated with higher accuracy levels than one or more unselected temporal boundaries from among the first set of temporal boundaries and the second set of temporal boundaries.” Emphasis Ours. Zhang renders obvious the claimed “higher accuracy levels” wherein, as discussed above, Zhang teaches utilizing machine learning, comprising a machine learning model can include a bi-directional gradient recurring unit and fully connected neural network layers, to generate a target set of temporal boundaries wherein the selected target set of temporal boundaries were selected based on scores and rankings that relate to being (see Zhang para 23, 32, 41 - breakpoint management system 160 implements a combination of machine learning models and search techniques to process and analyze each media item to determine a set of breakpoints that limit the disruption to playback of the media item). Zhang teaches determining the optimal time to present particular breakpoints in the media based on scores and ranking and which corresponds to “higher accuracy levels than one or more unselected temporal boundaries from among the first set of temporal boundaries and the second set of temporal boundaries” because the chose breakpoints are a more accurate position to optimally present a breakpoint for presentation of content. Zhang further teaches that the set of breakpoints chosen are stored in a media database as a list that can be accessed using information regarding the corresponding media item (para 67 wherein a report given its broadest reasonable interpretation comprises a description of an event, wherein in the particular case, the event is a breakpoint associated with a particular media). Therefore, it would have been obvious to one having ordinary skill in the art before the time of the applicant’s effective filing date to modify Mao identifying boundary points, the metadata-defined boundaries, and the I-frame-identified boundaries using machine learning by further incorporating known elements of Zhang’s invention for iteratively selecting subsets of target boundaries in order to improve the accuracy of content boundary detection through use of an iterative process that leverages machine learning. Regarding claim 2, “wherein the first set of temporal boundaries are associated with at least one of a CV/ML device or an instantaneous decoder refresh (IDR) frames placing encoder algorithm” is further rejected on obviousness grounds as discussed in the rejection of claim 1 wherein Mao para [0040] in step 402 of figure 4, determine scene boundaries of media content 300 by associating them with I frame positions. Regarding claim 3, “wherein the CV/ML temporal boundary report log includes information associated with automated generation of boundaries by the CV/ML device, the automated generation being utilized by the CV/ML device to infer CV/ML generated temporal boundaries utilizing an encode of the media content, the media content being analyzed by the encode for instantaneous decoder refresh (IDR) frame placement, the IDR frame placement being utilized to identify IDR frames and non-IDR frames associated with a third portion of the media content, individual ones of the IDR frames being followed by at least one of the non-IDR frames” is further rejected on obviousness grounds as discussed in the rejection of claims 1-2 wherein Mao teaches all the elements of the claim except a report as claimed wherein para [0037-0045] identifying boundaries of using metadata included with the video wherein number of encoders available to encode scenes and which encoding parameters may be used by specific encoders may be determined; Target resolutions and/or bit rates for subsequent transmission of scenes may be determined; the computing device may determine that each scene should be encoded; See para 33-36 regarding merging boundaries. Mao ([0044] discloses identifying visual elements using a machine learning algorithm. See also Zhang teaches generating a set of candidate breakpoints in a media item, and using a machine learning model to score the candidate breakpoints and iteratively select subsets of the candidate breakpoints to be a final set of breakpoints, which are stored as part of a bitstream. See para 11, 22, 29 regarding a list and figure 2. Zhang discloses a target set of boundaries in the form of a final set of breakpoints element 240 and 0-240 in figure 2. Regarding claim 8, “further comprising packaging the encoded media content as packaged media content” is further rejected on obviousness grounds as discussed in the rejection of claims 1-3 wherein Mao para [0030] discloses encoding visual elements with their associated audio content, and differentiating elements having audio content associated therewith (e.g. a newscaster), from silent elements (e.g. a stockticker).) Claim(s) 4-7, 13-16 are rejected under 35 U.S.C. 103 as being unpatentable over Mao; Weidong et al. US 20200288149 A1 (hereafter Mao) and in further view of Zhang; Wenbo et al. US 20210390130 A1 (hereafter Zhang) and in further view of Effinger; Charles et al. US 10841666 B1 (hereafter Effinger) and in further view of Oyman; Ozgur US 20160165185 A1 (hereafter Oyman). Regarding claim 4, “further comprising: determining a target temporal boundary report that includes the target set of temporal boundaries; and determining a default temporal boundary report that includes the second set of temporal boundaries and a default boundary report log associated with an encode of the media content” is further rejected on obviousness grounds as discussed in the rejection of claims 1-3 wherein Mao do not use the terms target and default with respect to the reports but Zhang teaches generating a set of candidate breakpoints in a media item, and using a machine learning model to score the candidate breakpoints and iteratively select subsets of the candidate breakpoints to be a final set of breakpoints, which are stored as part of a bitstream. See para 11, 22, 29 regarding a list and figure 2. Zhang discloses a target set of boundaries in the form of a final set of breakpoints element 240 and 0-240 in figure 2. See also Zhang (pars. [0022], [0023]) disclosing methods to score the result provided by ML-based boundary determination algorithms. Zhang further teaches that the set of breakpoints chosen are stored in a media database as a list that can be accessed using information regarding the corresponding media item (para 66-67 after several iterations, a default final set of breakpoints; see also wherein a report given its broadest reasonable interpretation comprises a description of an event, wherein in the particular case, the event is a breakpoint associated with a particular media). In an analogous art, Effinger teaches (col. 5, lines 52-64; col. 7, lines 40-54) (pars. [0022], [0023]), disclosing methods to score the result provided by ML-based boundary determination algorithms. In an analogous art, Oyman figure 3 discloses “SOP offer indicating: arbitrary and/or predefined ROI signaling support; actual transmitted ROI signaling; a description of each offered predefined ROI.” Therefore, it would have been obvious to one having ordinary skill in the art before the time of the applicant’s effective filing date to modify Mao and Zhang for identifying boundary points, the metadata-defined boundaries, and the I-frame-identified boundaries using machine learning and for iteratively selecting subsets of target boundaries in order to improve the accuracy of content boundary detection through use of an iterative process that leverages machine learning and/or analyze the media content (e.g., using machine learning and/or one or more graphics processing algorithms) to determine scene boundaries of the media content wherein a “default report” is generated from metadata by further incorporating known elements of Effinger’s and Oyman for the sender/client feedback information for ROI definition disclosed in Oyman into the scene classification and encoding system of Mao in view of Zhang, in order to improve efficiency of bandwidth usage by transmitting only those portions (ROI) of a video frame that have been requested by the client. Regarding claim 5, “further comprising: utilizing a second temporal boundary generation process on the media content to generate the default temporal boundary report; and utilizing a third temporal boundary generation process on the media content to generate the target temporal boundary report, the target temporal boundary report being a combination of the CV/ML temporal boundary report and the default boundary report” is further rejected on obviousness grounds as discussed in the rejection of claims 1-4 wherein Zhang teaches generating a set of candidate breakpoints in a media item, and using a machine learning model to score the candidate breakpoints and iteratively select subsets of the candidate breakpoints to be a final set of breakpoints, which are stored as part of a bitstream. See para 11, 22, 29 regarding a list and figure 2. Zhang discloses a target set of boundaries in the form of a final set of breakpoints element 240 and 0-240 in figure 2. See also Effinger teaches (col. 5, lines 52-64; col. 7, lines 40-54) (pars. [0022], [0023]), disclosing methods to score the result provided by ML-based boundary determination algorithms. See also Oyman figure 3 discloses “SOP offer indicating: arbitrary and/or predefined ROI signaling support; actual transmitted ROI signaling; a description of each offered predefined ROI.” Regarding claim 6, “wherein the target temporal boundary report has a higher level of accuracy than at least one of the CV/ML boundary report or the default temporal boundary report” is further rejected on obviousness grounds as discussed in the rejection of claims 1-5 wherein Effinger (col. 5, lines 52-64; col. 7 lines 40-54) and Zhang (abstract; pars. [0022], [0023]), where the terms higher score is interpreted as the content insertion point to a higher level of accuracy. Zhang further teaches that the set of breakpoints chosen are stored in a media database as a list that can be accessed using information regarding the corresponding media item (para 66-67 after several iterations, a default final set of breakpoints; see also wherein a report given its broadest reasonable interpretation comprises a description of an event, wherein in the particular case, the event is a breakpoint associated with a particular media). Regarding claim 7, “wherein encoding the media content is based at least in part on the target temporal boundary report and the default boundary report” is further rejected on obviousness grounds as discussed in the rejection of claims 1-6 wherein Effinger (col. 5, lines 52-64; col. 7 lines 40-54) and Zhang (abstract; pars. [0022], [0023]), where the terms higher score is interpreted as the content insertion point to a higher level of accuracy and Oyman figure 3 discloses “SOP offer indicating: arbitrary and/or predefined ROI signaling support; actual transmitted ROI signaling; a description of each offered predefined ROI.” Claim(s) 9, 19 are rejected under 35 U.S.C. 103 as being unpatentable over Mao; Weidong et al. US 20200288149 A1 (hereafter Mao) and in further view of Zhang; Wenbo et al. US 20210390130 A1 (hereafter Zhang) and in further view of Wu; Yongjun et al. US 10951960B2 (hereafter Wu). Regarding claim 9, “further comprising: generating a manifest link associated with the packaged media content; transmitting the manifest link to a destination device; receiving, from the destination device, a client request indicating a selection of the manifest link based at least in part on user input received via the destination device; determining that the client request is to stream the packaged media content; and causing streaming of the packaged media content via the destination device” wherein Mao teaches an invention comprising video encodingusing MPEG (para 58, 68) the invention does not use the term manifest but a person of ordinary skill in the art would have understood that manifests are a typical component of MPEG video transmission utilizing URL links. In an analogous art, Wu teaches the deficiency of Mao and Zhang (col. 2:52-67 to col. 4:1-52 - retrieval of these fragments can be much later, in the context of a VOD access, or nearly instantaneous (near real-time), in the context of streaming. In some implementations, a corresponding index file, commonly referred to as a manifest, may be provided to help retrieve these fragments. For example, the manifest may contain fragment data, such as a pointer, link (e.g., URL), redirect, or timing data associated with each fragment that ultimately allows an end user device, such as a media player, to receive each fragment at the appropriate time. The manifest may also include fragment sequence data, as well as various media data, such as a computer program or instruction for encoding or decoding the fragments (commonly known as a CODEC), a fragment title, etc.). Therefore, it would have been obvious to one having ordinary skill in the art before the time of the applicant’s effective filing date to modify Mao and Zhang for identifying boundary points, the metadata-defined boundaries, and the I-frame-identified boundaries using machine learning and for iteratively selecting subsets of target boundaries in order to improve the accuracy of content boundary detection through use of an iterative process that leverages machine learning and/or analyze the media content (e.g., using machine learning and/or one or more graphics processing algorithms) to determine scene boundaries of the media content wherein a “default report” is generated from metadata by further incorporating known elements of Wu for generating aa manifest to help a client device retrieve these fragments wherein the manifest may contain fragment data, such as a pointer, link (e.g., URL), redirect, or timing data associated with each fragment that ultimately allows an end user device, such as a media player, to receive each fragment at the appropriate time. Claim(s) 10, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Mao; Weidong et al. US 20200288149 A1 (hereafter Mao) and in further view of Zhang; Wenbo et al. US 20210390130 A1 (hereafter Zhang) and in further view of Effinger; Charles et al. US 10841666 B1 (hereafter Effinger) and in further view of Oyman; Ozgur US 20160165185 A1 (hereafter Oyman) and in further view of Kieft; Alexander J. et al. US 20200322401 A1 (hereafter Kieft). Regarding claim 10, “wherein encoding the media content comprises: encoding subtitles content as encoded subtitles content; and encoding thumbnails content as encoded thumbnails content, the encoded subtitles content and the encoded thumbnails content being temporally aligned with the encoded media content” Mao, Zhang, Effinger, Oyman are silent with respect to subtitles as claimed. However, Kieft discloses in an analogous art directed maintaining synchronicity among different elements of a live media stream, including synchronizing thumbnails and caption information. See [0017]. Therefore, it would have been obvious to one having ordinary skill in the art before the time of the applicant’s effective filing date to modify Mao, Zhang, Effinger, and Oyman for identifying boundary points, the metadata-defined boundaries, and the I-frame-identified boundaries using machine learning and for iteratively selecting subsets of target boundaries in order to improve the accuracy of content boundary detection through use of an iterative process that leverages machine learning and/or analyze the media content (e.g., using machine learning and/or one or more graphics processing algorithms) to determine scene boundaries of the media content by further incorporating known elements of Kieft to encode captions/subtitle content, as well as thumbnails, and to temporally align these elements with the rest of a media content stream, as disclosed in, and to incorporate these elements. Doing so would have entailed simply combining the prior art elements respectively disclosed in Mao, Zhang, and in Kieft, without changing their respective functions, and the combination would have yielded nothing more than predictable results for one of ordinary skill in the art. KSR Int'l Co. v. Teleflex Inc. See 2143.1.A. 550 U.S. at 416, 82 USPQ2d at 1395. CONCLUSION Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALFONSO CASTRO whose telephone number is (571)270-3950. The examiner can normally be reached on Monday to Friday from 10am to 6pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Nathan Flynn can be reached. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ALFONSO CASTRO/Primary Examiner, Art Unit 2421
Read full office action

Prosecution Timeline

Show 7 earlier events
Jun 26, 2026
Examiner Interview Summary
Jul 27, 2026
Response after Non-Final Action
Aug 11, 2026
Request for Continued Examination
Aug 15, 2026
Response after Non-Final Action
Aug 26, 2026
Non-Final Rejection mailed — §103
Sep 12, 2026
Interview Requested
Sep 25, 2026
Applicant Interview (Telephonic)
Sep 28, 2026
Examiner Interview Summary

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744966
CONTROL DEVICE, CONTROL METHOD, AND RECORDING MEDIUM
3y 0m to grant Granted Sep 22, 2026
Patent 12707120
LOW-LATENCY CONTENT DELIVERY OVER A PUBLIC NETWORK
2y 0m to grant Granted Aug 11, 2026
Patent 12689787
SYSTEM FOR PROGRAMMING CONTENT CHANNELS
2y 5m to grant Granted Jul 21, 2026
Patent 12647640
Information Processing Apparatus, Information Processing Method, and Program
5y 1m to grant Granted Jun 02, 2026
Patent 12641188
METHOD OF BROADCASTING REAL-TIME ON-LINE COMPETITIONS AND APPARATUS THEREFOR
2y 8m to grant Granted May 26, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
51%
Grant Probability
70%
With Interview (+19.4%)
3y 8m (~1y 8m remaining)
Median Time to Grant
High
PTA Risk
Based on 451 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month