Prosecution Insights
Last updated: October 02, 2026
Application No. 18/817,675

VIDEO CONTENT TO SLIDE GENERATION

Non-Final OA §103
Filed
Aug 28, 2024
Examiner
NGUYEN, CAO H
Art Unit
2171
Tech Center
2100 — Computer Architecture & Software
Assignee
Open Text Corporation
OA Round
1 (Non-Final)
91%
Grant Probability
Favorable
1-2
OA Rounds
5m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 91% — above average
91%
Career Allowance Rate
1048 granted / 1153 resolved
+35.9% vs TC avg
Moderate +7% lift
Without
With
+7.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
28 currently pending
Career history
1164
Total Applications
across all art units

Statute-Specific Performance

§101
9.7%
-30.3% vs TC avg
§103
47.5%
+7.5% vs TC avg
§102
14.4%
-25.6% vs TC avg
§112
5.1%
-34.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 1153 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-3, 5-8, 11-13, 15-17 and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang et al. (US Patent Application Publication No. 2018/0082124) in view of Parmar et al. (US Patent Application Publication No. 2021/0076105). Regarding claim 1, Zhang discloses a method for compressing a video presentation, comprising [see abstract, extracting a set of keyframes from a video presentation. A subset of the keyframes is selected to present to a user based on a user preference. Annotations are generated for the subset of the keyframes]: capturing a first still image [keyframes] from a first frame of the video presentation [see para. 0030-0033 and figure 3; extracting a set of keyframes from a video presentation, many different techniques may be appropriate for extracting keyframes from video presentations. Further details of extracting keyframes from video presentations]. writing the first still image to a document [see para. 0047-0048 and figure 2; generating personalized annotations. The personalized annotations may be generated for the personalized keyframes. The personalized annotations may be generated based on the personalization settings associated with the user, the personalized annotations may be generated from annotations obtained from other users, annotations recommended to other users, text extracted from the video presentation, audio content extracted from the video; which equivalent to writing the still images to a document]; capturing a second still image from a second frame of the video presentation [see para. 0046 and figure 5; Method includes extracting a set of capstone keyframes. The capstone keyframes extracted from a video presentation. The capstone keyframes extracted based on clear cuts within the video presentation. Additionally, the set of capstone keyframes extracted account for visual subject matter of the capstone keyframes]; comparing the first still image to the second still image to determine whether the difference is greater than a threshold [see para. 0022, 0040 and figure 5; The capstone keyframes from the video presentation may be frames just prior to a clear cut. This facilitate maximizing content presented to a user in the keyframes provided to the user, keyframes presented to users with references to portions of the video presentation with which the keyframes are associated. When capstone keyframes are presented to a user with such a reference, the reference cause the user to be directed to the beginning of the portion of the video from which the capstone keyframe is extracted. This portion of the video identified by, for example, detecting a previous clear cut, presentation content analysis (e.g., for when there are several clear cuts during the discussion of a single content slide); which corresponds to the description of extracting keyframes based on clear cuts]; and when the difference is greater than the threshold, and when the difference is not greater than the threshold [see para. 0019, 0039 and figures 1-3; color histograms may be made and compared for adjacent frames. When color histograms between two frames in a video presentation exceed a threshold, a clear cut may be considered to have occurred between the two frames. The threshold predefined or adaptive. An adaptive threshold based on a detected size of changes in histograms of frames over time, template subtraction techniques applied to detect clear cuts in a video presentation. Template subtraction involve subtracting contents of a frame from a prior frame to detect whether changes to the frame are local to one area of the frame, or change a large portion of the frame. Extracting the set of keyframes also includes detecting a clear cut. Clear cuts detected in a content area containing a presentation slide. Clear cuts detected using one of, histogram thresholds, template subtraction, grid division analysis which equivalent to capture still image from frames of a video presentation, comparing them a different threshold]; however; Zhang fails to explicitly teach writing the second still image to the document and omitting writing the second still image to the document. Parmar discloses writing the second still image to the document and omitting writing the second still image to the document [see para. 0007; wherein the video of the presentation includes images of presentation slides; and/or wherein the media stream further includes slide data from a slide presenting device; and/or further comprising: associating a time stamp metadata to elements of slide data; time ordering the elements of the slide data; displaying in the one or more display panes, elements of the slide data; and enabling a playback of the elements of the slide data, wherein the displayed elements of the slide data are time matched to at least one of a displayed elements of the segmented video and transcribed text, and when a different time selection is made by the user; and/or wherein a first pane of the one or more display panes is a view of the elements of the slide data and a second pane of the one or more display panes is a view of the elements of the transcribed text; and/or further comprising, concurrently displaying a plurality of elements of the segmented video in a video pane of the one or more display panes; which corresponds to convert a video of a slide presentation into an editable document by generating keyframes and still images of the presentation content and writing still images]. It would have been obvious to one of an ordinary skill in the art, having the teachings of Zhang and Parmar before the affective filing date of the claimed invention to modify, Zhang’s keyframes extraction process to include writing the selected still images into a document, as taught by Parmar. One would have been motivated to make such a combination in order to provide the result of a method that compresses the video presentation by writing the still images that exceeds a difference threshold into a document. Regarding claims 2 and 12, Zhang discloses wherein comparing the first still image to the second still image comprises: applying a first mask to mask the first frame to exclude a first human subject image to produce a masked first still image; and applying a second mask to the second frame to exclude a second human subject image to produce a masked second still image [see para. 0035-0038 and figure 3; Extracting the set of keyframes includes avoiding selecting as keyframes, video frames featuring a person. This may be done by removing as potential keyframes, frames of the video presentation that feature a talking person. These frames may be identified by, for example, performing facial recognition on keyframes. Frames featuring a talking person may make inferior keyframes for video presentations because they do not contain pictographic content related to the subject matter of the video presentation]. Parmar discloses wherein comparing the first still image to the second still image to determine if the difference is greater than the threshold comprises comparing the masked first still image to the masked second still image to determine whether the difference between the first masked image and the second masked image is greater than the threshold [see para. 0256-0260; a pixelwise segmentation mask, or as polygonal outlines. f. Identifies whether the surface is chalkboard, whiteboard, glassboard, smartboard, paper surface, or other writable material. Operations from Person Detector (Extract and/or Mask) Module: a. People are the most common distractors in front of writing surfaces, so the exemplary system is able implement a dedicated detector to detect them (so as distractors they can be ignored by algorithms focusing on writing). b. The algorithm is aware and also learns what a human is and generates a pixelwise mask (each pixel is assigned a probability of “person” vs “non-person”), polygonal outline, and/or pose skeleton]. One would have been motivated to make such a combination in order to provide the result of a method that compresses the video presentation by writing the still images that exceeds a difference threshold into a document. Regarding claims 3 and 13, Parmar discloses further comprising: upon determining the ratio of the second mask to the second still image is greater than a mask threshold, determining the difference between the masked first still image and the masked second still image is not greater than the threshold [see para. 0380 and figure 5B; Recognizing the duration of writing is described in the section on writing change detection. [0381] i. For example: save the last N sampled video frames. For each pixel: if the human mask blocks most of the N frames, then don't update that pixel (it will thus remain inpainted with whatever was there before the person walked in front); otherwise update it with the average of the non-masked pixels]. Regarding claims 5 and 15, Zhang discloses further comprising: calculating a first image hash of the first still image; and calculating a second image hash of the second still image; and wherein comparing the first still image to the second still image to determine whether the difference is greater than the threshold comprises comparing the first image hash to the second image hash to determine whether the difference is greater than the threshold [see para. 0019; color histograms may be made and compared for adjacent frames. When color histograms between two frames in a video presentation exceed a threshold, a clear cut may be considered to have occurred between the two frames. The threshold may be predefined or adaptive. An adaptive threshold may be based on a detected size of changes in histograms of frames over time. In another example, template subtraction techniques may be applied to detect clear cuts in a video presentation. Template subtraction may involve subtracting contents of a frame from a prior frame to detect whether changes to the frame are local to one area of the frame, or change a large portion of the frame]. Regarding claims 6 and 16, Zhang discloses wherein calculating the first image hash of the first still image and calculating the second image hash of the second still image each comprise calculating the hash further determined from at least one of image average analysis, perceptual analysis, difference analysis or wavelet analysis on structures within each of the first still image and the second still image [see para. 0039, Extracting the set of keyframes also includes detecting a clear cut. Clear cuts may be detected in a content area containing a presentation slide. Clear cuts may be detected using one of, histogram thresholds, template subtraction, grid division analysis, and so forth. Detecting clear cuts may facilitate identifying when presentation slides have advanced to a new slide. Because frames just prior to clear cuts may have more content than a new slide, the frames preceding clear cuts may be prioritized as keyframes]. Regarding claims 7 and 17, Zhang discloses further comprising: creating a first grayscale image of the first still image; and creating a second grayscale image of the second still image; and wherein comparing the first still image to the second still image to determine whether the difference is greater than the threshold comprises comparing the first grayscale image to the second grayscale image to determine whether the difference is greater than the threshold [see para. 0023 and figure 1; performing facial recognition on capstone keyframes may not be an effective method for certain types of video presentations. For example, frame illustrates a frame of a video presentation having several different content areas. In frame, presentation slides are on the left, a video feed of the presenter is in the lower right, and a blank content area is in the upper right. Other configurations of content areas are also possible]. Regarding claim 8, Zhang discloses wherein the first still image and the second still image are separated by at least one intervening frame of the video presentation [see para. 0026; Once keyframes personalized for a user, annotations to be provided to the user along with the keyframes may also be generated. As with the specific keyframes provided, the annotations provided may also be personalized to that user (e.g., based on user settings, based on past user behavior, user interactions). The actual annotations provided to the user provided from a variety of sources. For example, the annotations generated from notes generated by the user, notes generated and/or recommended by other users, slide content, speech to text of content from the video presentation]. Regarding claims 11 and 20, Zhang discloses system for compressing a video presentation, comprising: at processor coupled to a computer memory comprising computer-readable instructions to cause the processor to [see para. 0029 and figure 2; a non-transitory computer-readable medium storing computer-executable instructions. The instructions, when executed by a computer, may cause the computer to perform method]: capture a first still image [keyframes] from a first frame of the video presentation [see para. 0030-0033 and figure 3; extracting a set of keyframes from a video presentation, many different techniques may be appropriate for extracting keyframes from video presentations. Further details of extracting keyframes from video presentations]. write the first still image to a document [see para. 0047-0048 and figure 2; generating personalized annotations. The personalized annotations may be generated for the personalized keyframes. The personalized annotations may be generated based on the personalization settings associated with the user, the personalized annotations may be generated from annotations obtained from other users, annotations recommended to other users, text extracted from the video presentation, audio content extracted from the video; which equivalent to writing the still images to a document]; capture a second still image from a second frame of the video presentation [see para. 0046 and figure 5; Method includes extracting a set of capstone keyframes. The capstone keyframes extracted from a video presentation. The capstone keyframes extracted based on clear cuts within the video presentation. Additionally, the set of capstone keyframes extracted account for visual subject matter of the capstone keyframes]; compare the first still image to the second still image to determine whether the difference is greater than a threshold [see para. 0022, 0040 and figure 5; The capstone keyframes from the video presentation may be frames just prior to a clear cut. This facilitate maximizing content presented to a user in the keyframes provided to the user, keyframes presented to users with references to portions of the video presentation with which the keyframes are associated. When capstone keyframes are presented to a user with such a reference, the reference cause the user to be directed to the beginning of the portion of the video from which the capstone keyframe is extracted. This portion of the video identified by, for example, detecting a previous clear cut, presentation content analysis (e.g., for when there are several clear cuts during the discussion of a single content slide); which corresponds to the description of extracting keyframes based on clear cuts]; and when the difference is greater than the threshold, and when the difference is not greater than the threshold [see para. 0019, 0039 and figures 1-3; color histograms may be made and compared for adjacent frames. When color histograms between two frames in a video presentation exceed a threshold, a clear cut may be considered to have occurred between the two frames. The threshold predefined or adaptive. An adaptive threshold based on a detected size of changes in histograms of frames over time, template subtraction techniques applied to detect clear cuts in a video presentation. Template subtraction involve subtracting contents of a frame from a prior frame to detect whether changes to the frame are local to one area of the frame, or change a large portion of the frame. Extracting the set of keyframes also includes detecting a clear cut. Clear cuts detected in a content area containing a presentation slide. Clear cuts detected using one of, histogram thresholds, template subtraction, grid division analysis which equivalent to capture still image from frames of a video presentation, comparing them a different threshold]; however; Zhang fails to explicitly teach write the second still image to the document and omitting writing the second still image to the document; and cause a data storage component to store the document. Parmar discloses write the second still image to the document and omitting writing the second still image to the document; and cause a data storage component to store the document [see para. 0007; wherein the video of the presentation includes images of presentation slides; and/or wherein the media stream further includes slide data from a slide presenting device; and/or further comprising: associating a time stamp metadata to elements of slide data; time ordering the elements of the slide data; displaying in the one or more display panes, elements of the slide data; and enabling a playback of the elements of the slide data, wherein the displayed elements of the slide data are time matched to at least one of a displayed elements of the segmented video and transcribed text, and when a different time selection is made by the user; and/or wherein a first pane of the one or more display panes is a view of the elements of the slide data and a second pane of the one or more display panes is a view of the elements of the transcribed text; and/or further comprising, concurrently displaying a plurality of elements of the segmented video in a video pane of the one or more display panes; which corresponds to convert a video of a slide presentation into an editable document by generating keyframes and still images of the presentation content and writing still images]. It would have been obvious to one of an ordinary skill in the art, having the teachings of Zhang and Parmar before the affective filing date of the claimed invention to modify, Zhang’s keyframes extraction process to include writing the selected still images into a document, as taught by Parmar. One would have been motivated to make such a combination in order to provide the result of a method that compresses the video presentation by writing the still images that exceeds a difference threshold into a document. Allowable Subject Matter Claims 4, 9, 10 , 14 , 18 and 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure (See PTO-892). Adock et al. (US 2011/0081075) discloses a system and method for identifying key frames of a presentation video that include stationary informational content. A sequence of frames is obtained from a presentation video and differences of pixel values between consecutive frames of the sequence of frames are computed. Sets of consecutive frames that are stationary are identified, wherein consecutive frames that are stationary have a proportion of changed pixel values below a first predetermined threshold, and wherein pixel values are deemed to be changed when the difference between the pixel values for corresponding pixels in consecutive frames exceeds a second predetermined threshold. Bovik et al. (US 2018/0268864) discloses a video frame containing a whiteboard image is converted into a black and white image for the detection of boundaries. These boundaries are classified as horizontal or vertical lines. Quadrangles are then formed using spatial arrangements of these lines. The quadrangles that are most likely to spatially coincide with the boundaries of the whiteboard image are identified. The quadrangles are then sorted (ranked) based on specific characteristics, such as size and position. The area corresponding to the identified quadrangle in the video frame is then cropped. Furthermore, the speaker in the video frame can be removed based on detecting changes that are characteristic of movements of a speaker. In this manner, the visual and educational experience involved in the recording of classroom lectures or other such presentation is improved. A reference to specific paragraphs, columns, pages, or figures in a cited prior art reference is not limited to preferred embodiments or any specific examples. It is well settled that a prior art reference, in its entirety, must be considered for all that it expressly teaches and fairly suggests to one having ordinary skill in the art. Stated differently, a prior art disclosure reading on a limitation of Applicant's claim cannot be ignored on the ground that other embodiments disclosed were instead cited. Therefore, the Examiner's citation to a specific portion of a single prior art reference is not intended to exclusively dictate, but rather, to demonstrate an exemplary disclosure commensurate with the specific limitations being addressed. In re Heck, 699 F.2d 1331, 1332-33,216 USPQ 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006,1009, 158 USPQ 275, 277 (CCPA 1968)). In re: Upsher-Smith Labs. v. Pamlab, LLC, 412 F.3d 1319, 1323, 75 USPQ2d 1213, 1215 (Fed. Cir. 2005); In re Fritch, 972 F.2d 1260, 1264, 23 USPQ2d 1780, 1782 (Fed. Cir. 1992); Merck & Co. v. Biocraft Labs., Inc., 874 F.2d 804, 807, 10 USPQ2d 1843, 1846 (Fed. Cir. 1989); In re Fracalossi, 681 F.2d 792,794 n.1,215 USPQ 569, 570 n.1 (CCPA 1982); In re Lamberti, 545 F.2d 747, 750, 192 USPQ 278, 280 (CCPA 1976); In re Bozek, 416 F.2d 1385, 1390, 163 USPQ 545, 549 (CCPA 1969). Any inquiry concerning this communication or earlier communications from the examiner should be directed to CAO H NGUYEN whose telephone number is (571)272-4053. The examiner can normally be reached on Mon-Fri 9am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kieu Vu can be reached on 571-272-4057. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CAO H NGUYEN/Primary Examiner, Art Unit 2171
Read full office action

Prosecution Timeline

Aug 28, 2024
Application Filed
Aug 03, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748568
MULTI-MODAL INTERFACES
2y 5m to grant Granted Sep 29, 2026
Patent 12739218
METHODS, APPARATUSES, SYSTEMS AND STORAGE MEDIA FOR PROCESSING A LINK IN A CONVERSATION
2y 8m to grant Granted Sep 15, 2026
Patent 12724524
CURSOR PROMPT INTERFACES FOR FACILITATING MULTIPLE FUNCTIONALITY
2y 7m to grant Granted Sep 01, 2026
Patent 12724627
Message Notification Method and Apparatus
2y 5m to grant Granted Sep 01, 2026
Patent 12717467
INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND INFORMATION PROCESSING PROGRAM
2y 5m to grant Granted Aug 25, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
91%
Grant Probability
98%
With Interview (+7.4%)
2y 6m (~5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 1153 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month