Prosecution Insights
Last updated: August 18, 2026
Application No. 18/597,438

METADATA-BASED THUMBNAIL IMAGE GENERATION FOR PRESENTATION ON A CONTENT PLATFORM

Non-Final OA §103
Filed
Mar 06, 2024
Examiner
CHIN, MICHELLE
Art Unit
2614
Tech Center
2600 — Communications
Assignee
Google LLC
OA Round
3 (Non-Final)
86%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 86% — above average
86%
Career Allowance Rate
556 granted / 650 resolved
+23.5% vs TC avg
Moderate +11% lift
Without
With
+11.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 2m
Avg Prosecution
24 currently pending
Career history
674
Total Applications
across all art units

Statute-Specific Performance

§101
9.7%
-30.3% vs TC avg
§103
70.4%
+30.4% vs TC avg
§102
5.7%
-34.3% vs TC avg
§112
1.8%
-38.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 650 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status 1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 2. A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/01/2026 has been entered. Response to Amendment 3. Acknowledgement is made of amendment filed on June 01, 2026, in which claims 1, 8, 17 and 23 are amended, and claims 1-28 are still pending. Response to Arguments 4. Applicant's arguments, filed on June 01, 2026, with respect to Claims 1-28 have been fully considered and they are not persuasive. 5. With regards to arguments for independent claims 1, 8, 17 and 23, applicants argue that Willis et al. (US 2023/0104396), O’Neill (US 2024/0412542 A1) and Rafati et al. (US 2014/0099034 A1) fail to disclose automatically generating a textual prompt from the one or more metadata items characterizing the respective one or more expressive aspects associated with the collection of media items, the textual prompt describing visual characteristics of the thumbnail image to be generated; causing an artificial intelligence (AI) generative model to process the textual prompt describing the visual characteristics of the thumbnail image; and obtaining one or more outputs from the AI generative model. The examiner respectfully agrees and moots in view of the new grounds of rejections regarding claims 1, 8, 17 and 23, since in Cragg et al. (US 2024/0273670 A1) teaches (“A prompt generation network is trained to infer a prompt based on the image and an image generation network generates an expanded image based on the prompt. For example, the prompt can be generated based on metadata (e.g., location, time, color, lighting information) associated with the image. The prompt is then fed to an image generation network (e.g., a diffusion model) to generate an expanded image that includes an outpainted region. Accordingly, users can easily expand the crop of an image to improve the image composition, change the aspect ratio, or generate new and creative variations of the same image.” [0003] “Image processing apparatus 110 includes a computer implemented network comprising a machine learning model. The machine learning model includes a prompt generation network and a diffusion model.” [0038] “machine learning model 225 obtains an image and a target dimension for expanding the image. In some examples, machine learning model 225 identifies an expanded region having the target dimension, where the expanded region includes the image and the outpainted region.” [0052] “The text prompt 335 can be encoded using a text encoder 340 (e.g., a multi-modal encoder) to obtain guidance features 345 in guidance space 350. The guidance features 345 can be combined with the noisy images 320 at one or more layers of the reverse diffusion process 325 to ensure that the output image 330 includes content described by the text prompt 335.” [0067] “The image processing apparatus generates an expanded image in response to the input (e.g., the image and the target dimension) and displays the expanded image to the user. The expanded image includes new pixels that are generated by the image processing apparatus in an outpainted region. The outpainted region is a region of difference between the image (e.g., input image) and the target dimension. The content of the pixels of the expanded image that corresponds to the outpainted region is generated by the image processing apparatus based on an inferred prompt. According to some embodiments, the image processing apparatus automatically infers the prompt based on the metadata of the image.” [0084] “the image cropping interface includes interface element(s) 815. Interface element 815 shows a set of candidate expanded image 800 (i.e., low resolution images or thumbnails) that may be different in terms of style, size, shading effect, color, etc. For example, a user selects a thumbnail from the set of candidate expanded images via interface element 815 to obtain expanded image 800.” [0106]) Cragg teaches the prompt generated based on metadata, the prompt is then fed to an image generation network to generate an expanded image, the expanded image or thumbnail can describe style, size, shading effect, color, etc. Therefore, Cragg teaches the arguments of the limitations for claims 1, 8, 17 and 23 as it is recited. Claim Rejections - 35 USC § 103 6. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 7. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 8. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. 9. Claim(s) 1, 2, 5, 7-9, 15,17, 18, 21, 23 and 24 is/are rejected under 35 U.S.C. 103 as being unpatentable over Willis et al. (US 2023/0104396) in view of Cragg et al. (US 2024/0273670 A1). 10. With reference to claim 1, Willis teaches A method, comprising: receiving, by a processing device, a request initiated by a user to generate a thumbnail image to be associated with a collection of media items stored by a content platform; (“The method 600 includes authenticating at 602 a user and enabling the user to log in to an administrator portal of an account supported by the media server. The method 600 includes receiving data at 604 uploaded by the user and storing the data on the image server for cloud-based access. The method 600 includes providing at 606 a user interface to the user that enables the user to organize and manage the data collection. The method 600 includes receiving at 608 a thumbnail selection approval from the user, wherein the thumbnail comprises a screen capture from one or more data objects within the data collection. The method 600 includes generating at 610 a unique code that provides instructions to a personal device to access the data collection. The method 600 includes merging at 612 the unique code and the thumbnail selection to generate a scannable thumbnail.” [0063]) Willis also teaches identifying one or more metadata items characterizing respective one or more expressive aspects associated with the collection of media items; (“The data uploaded by the user may include one or more independent data objects such as videos, images, written works, numerical data, historical data, and so forth. The one or more data objects may be grouped together to generate a data collection. The data collection may be grouped together by the neural network 114 as described herein and/or manually grouped together by the user. One data collection may include one or more data objects that are related based on geographical location, time, quality, subject matter, or some other metric. The data collection may include one or more data objects that are stitched together on the media library 112 using common metadata to indicate that each of the one or more data objects should be associated with the same data collection.” [0064] “The user may submit tags to be associated with the video 806, and these tags represent metadata for the video 806. The tags may, for example, who is depicted in the video, the subject of the video, the time or season the video was captured, and so forth. The tags will be stored in association with the video 806 on the media library 112. … The combination of the unique code 810 and the thumbnail 808 produces a scannable thumbnail. The scannable thumbnail may be captured or scanned with a sensor and provide instructions to redirect to a website, file system, database, and so forth. The unique code 810 may provide instructions to access a website where the video 806 may be viewed. … The thumbnail 808 may include an image frame from a video 806, a screen capture, an image, a document, a graphic, and so forth. The unique code 810 may redirect a computing device to access additional media associated with the thumbnail 808.” [0074-0076]) Willis further teaches the one or more outputs specifying respective one or more thumbnail images. (“The media classification component 304 prioritizes data objects based at least in part on outputs from the neural network 114.” [0045] “The thumbnail component 312 may include a neural network for selecting an image or screen capture as described further herein. The thumbnail component 312 is further configured to generate a scannable thumbnail as described herein by merging a unique code (such as a QR code) with a selected thumbnail.” [0049-0050]) PNG media_image1.png 681 575 media_image1.png Greyscale Willis does not explicitly teach automatically generating a textual prompt from the one or more metadata items characterizing the respective one or more expressive aspects associated with the collection of media items, the textual prompt describing visual characteristics of the thumbnail image to be generated; causing an artificial intelligence (AI) generative model to process the textual prompt describing the visual characteristics of the thumbnail image; and obtaining one or more outputs from the AI generative model. This is what Cragg teaches (“A prompt generation network is trained to infer a prompt based on the image and an image generation network generates an expanded image based on the prompt. For example, the prompt can be generated based on metadata (e.g., location, time, color, lighting information) associated with the image. The prompt is then fed to an image generation network (e.g., a diffusion model) to generate an expanded image that includes an outpainted region. Accordingly, users can easily expand the crop of an image to improve the image composition, change the aspect ratio, or generate new and creative variations of the same image.” [0003] “Image processing apparatus 110 includes a computer implemented network comprising a machine learning model. The machine learning model includes a prompt generation network and a diffusion model.” [0038] “machine learning model 225 obtains an image and a target dimension for expanding the image. In some examples, machine learning model 225 identifies an expanded region having the target dimension, where the expanded region includes the image and the outpainted region.” [0052] “The text prompt 335 can be encoded using a text encoder 340 (e.g., a multi-modal encoder) to obtain guidance features 345 in guidance space 350. The guidance features 345 can be combined with the noisy images 320 at one or more layers of the reverse diffusion process 325 to ensure that the output image 330 includes content described by the text prompt 335.” [0067] “The image processing apparatus generates an expanded image in response to the input (e.g., the image and the target dimension) and displays the expanded image to the user. The expanded image includes new pixels that are generated by the image processing apparatus in an outpainted region. The outpainted region is a region of difference between the image (e.g., input image) and the target dimension. The content of the pixels of the expanded image that corresponds to the outpainted region is generated by the image processing apparatus based on an inferred prompt. According to some embodiments, the image processing apparatus automatically infers the prompt based on the metadata of the image.” [0084] “the image cropping interface includes interface element(s) 815. Interface element 815 shows a set of candidate expanded image 800 (i.e., low resolution images or thumbnails) that may be different in terms of style, size, shading effect, color, etc. For example, a user selects a thumbnail from the set of candidate expanded images via interface element 815 to obtain expanded image 800.” [0106]) Cragg teaches the prompt generated based on metadata, the prompt is then fed to an image generation network to generate an expanded image, the expanded image or thumbnail can describe style, size, shading effect, color, etc. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Cragg into Willis, in order to increase quality and sharpness of the generated images. 11. With reference to claim 2, Willis teaches the one or more expressive aspects associated with the collection of media items comprise at least one of: a genre associated with the collection of media items, a mood associated with the collection of media items, an emotion associated with the collection of media items, a lyrics associated with the collection of media items, a rhythm associated with the collection of media items, an instrumentation associated with the collection of media items, a vocal style associated with the collection of media items, a production style associated with the collection of media items, a cultural context associated with the collection of media items, or a theme associated with the collection of media items. (“The data uploaded by the user may include one or more independent data objects such as videos, images, written works, numerical data, historical data, and so forth. The one or more data objects may be grouped together to generate a data collection.” [0064] “FIG. 8A is a screenshot of the user interface 800 illustrating a page for uploading files to the media platform 102. The files may include media to be stored on the media server 104 and/or media library 112. The user may upload files by connecting with an external URL (Uniform Resource Locator) at 804. The user may additionally upload files that are stored locally on the user's computer or remotely on a cloud-based storage solution. The user may upload any suitable file type, including, for example, text files, video files, image files, music files, and specialty file types that may be associated with a certain program or application.” [0069]) 12. With reference to claim 5, Willis teaches receiving, via a user interface (UI), an input identifying a chosen thumbnail image of the one or more thumbnail images; and associating the chosen thumbnail image with the collection of media items. (“The method 600 includes authenticating at 602 a user and enabling the user to log in to an administrator portal of an account supported by the media server. The method 600 includes receiving data at 604 uploaded by the user and storing the data on the image server for cloud-based access. The method 600 includes providing at 606 a user interface to the user that enables the user to organize and manage the data collection. The method 600 includes receiving at 608 a thumbnail selection approval from the user, wherein the thumbnail comprises a screen capture from one or more data objects within the data collection. The method 600 includes generating at 610 a unique code that provides instructions to a personal device to access the data collection. The method 600 includes merging at 612 the unique code and the thumbnail selection to generate a scannable thumbnail. The data uploaded by the user may include one or more independent data objects such as videos, images, written works, numerical data, historical data, and so forth. The one or more data objects may be grouped together to generate a data collection.” [0063-0064]) 13. With reference to claim 7, Willis teaches causing the collection of media items to be presented in a first display area of the UI; and causing the chosen thumbnail to be presented in a second display area of the UI, wherein the second display area of the Ul is presented above the first display area of the UI. (“The method 600 includes authenticating at 602 a user and enabling the user to log in to an administrator portal of an account supported by the media server. The method 600 includes receiving data at 604 uploaded by the user and storing the data on the image server for cloud-based access. The method 600 includes providing at 606 a user interface to the user that enables the user to organize and manage the data collection. The method 600 includes receiving at 608 a thumbnail selection approval from the user, wherein the thumbnail comprises a screen capture from one or more data objects within the data collection. The method 600 includes generating at 610 a unique code that provides instructions to a personal device to access the data collection. The method 600 includes merging at 612 the unique code and the thumbnail selection to generate a scannable thumbnail. The data uploaded by the user may include one or more independent data objects such as videos, images, written works, numerical data, historical data, and so forth. The one or more data objects may be grouped together to generate a data collection.” [0063-0064] “FIG. 8F illustrates wherein the unique code 810 is displayed in the center of the thumbnail 808. The combination of the unique code 810 and the thumbnail 808 produces a scannable thumbnail. The scannable thumbnail may be captured or scanned with a sensor and provide instructions to redirect to a website, file system, database, and so forth. The unique code 810 may provide instructions to access a website where the video 806 may be viewed.” [0075]) 14. Claim 8 is similar in scope to claim 1, and thus is rejected under similar rationale. Willis additionally teaches identifying one or more stylistic features specified by the user; (“The method 600 includes receiving data at 604 uploaded by the user and storing the data on the image server for cloud-based access. The method 600 includes providing at 606 a user interface to the user that enables the user to organize and manage the data collection. The method 600 includes receiving at 608 a thumbnail selection approval from the user, wherein the thumbnail comprises a screen capture from one or more data objects within the data collection. The method 600 includes generating at 610 a unique code that provides instructions to a personal device to access the data collection. … The data uploaded by the user may include one or more independent data objects such as videos, images, written works, numerical data, historical data, and so forth. The one or more data objects may be grouped together to generate a data collection.” [0063-0064] “FIG. 8A is a screenshot of the user interface 800 illustrating a page for uploading files to the media platform 102. The files may include media to be stored on the media server 104 and/or media library 112. The user may upload files by connecting with an external URL (Uniform Resource Locator) at 804. The user may additionally upload files that are stored locally on the user's computer or remotely on a cloud-based storage solution. The user may upload any suitable file type, including, for example, text files, video files, image files, music files, and specialty file types that may be associated with a certain program or application.” [0069] “the user has uploaded a video 806 using the drag and drop 802 box. The user may continue to navigate through the process by clicking “next.” The user may play the video 806 to identify an image frame to use as a thumbnail to represent the video. The video 806 may be provided to a neural network trained to identify one or more optimal image frames that may be selected as the thumbnail.” [0071]) Willis does not explicitly teach automatically generating a textual prompt from the one or more stylistic features, the textual prompt describing visual characteristics of the thumbnail image to be generated; This is what Cragg teaches (“A prompt generation network is trained to infer a prompt based on the image and an image generation network generates an expanded image based on the prompt. For example, the prompt can be generated based on metadata (e.g., location, time, color, lighting information) associated with the image. The prompt is then fed to an image generation network (e.g., a diffusion model) to generate an expanded image that includes an outpainted region. Accordingly, users can easily expand the crop of an image to improve the image composition, change the aspect ratio, or generate new and creative variations of the same image.” [0003] “Image processing apparatus 110 includes a computer implemented network comprising a machine learning model. The machine learning model includes a prompt generation network and a diffusion model.” [0038] “machine learning model 225 obtains an image and a target dimension for expanding the image. In some examples, machine learning model 225 identifies an expanded region having the target dimension, where the expanded region includes the image and the outpainted region.” [0052] “The text prompt 335 can be encoded using a text encoder 340 (e.g., a multi-modal encoder) to obtain guidance features 345 in guidance space 350. The guidance features 345 can be combined with the noisy images 320 at one or more layers of the reverse diffusion process 325 to ensure that the output image 330 includes content described by the text prompt 335.” [0067] “The image processing apparatus generates an expanded image in response to the input (e.g., the image and the target dimension) and displays the expanded image to the user. The expanded image includes new pixels that are generated by the image processing apparatus in an outpainted region. The outpainted region is a region of difference between the image (e.g., input image) and the target dimension. The content of the pixels of the expanded image that corresponds to the outpainted region is generated by the image processing apparatus based on an inferred prompt. According to some embodiments, the image processing apparatus automatically infers the prompt based on the metadata of the image.” [0084] “the image cropping interface includes interface element(s) 815. Interface element 815 shows a set of candidate expanded image 800 (i.e., low resolution images or thumbnails) that may be different in terms of style, size, shading effect, color, etc. For example, a user selects a thumbnail from the set of candidate expanded images via interface element 815 to obtain expanded image 800.” [0106]) Cragg teaches the prompt generated based on metadata, the prompt is then fed to an image generation network to generate an expanded image, the expanded image or thumbnail can describe style, size, shading effect, color, etc. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Cragg into Willis, in order to increase quality and sharpness of the generated images. 15. With reference to claim 9, Willis teaches identifying the one or more stylistic features specified by the user further comprises: causing a list of stylistic features to be presented via a user interface (UI); (“The method 600 includes receiving data at 604 uploaded by the user and storing the data on the image server for cloud-based access. The method 600 includes providing at 606 a user interface to the user that enables the user to organize and manage the data collection. The method 600 includes receiving at 608 a thumbnail selection approval from the user, wherein the thumbnail comprises a screen capture from one or more data objects within the data collection. The method 600 includes generating at 610 a unique code that provides instructions to a personal device to access the data collection. … The data uploaded by the user may include one or more independent data objects such as videos, images, written works, numerical data, historical data, and so forth. The one or more data objects may be grouped together to generate a data collection.” [0063-0064] “FIG. 8A is a screenshot of the user interface 800 illustrating a page for uploading files to the media platform 102. The files may include media to be stored on the media server 104 and/or media library 112. The user may upload files by connecting with an external URL (Uniform Resource Locator) at 804. The user may additionally upload files that are stored locally on the user's computer or remotely on a cloud-based storage solution. The user may upload any suitable file type, including, for example, text files, video files, image files, music files, and specialty file types that may be associated with a certain program or application.” [0069] “Example 9 is a method as in any of Examples 1-8, further comprising generating a data collection playlist comprising the one or more data objects presented in a sequence as a video and publishing the data collection playlist on a webpage, wherein scanning the scannable thumbnail with a computing device provides instructions to the computing device to direct to the webpage to access the data collection playlist.” [0099]) Willis also teaches receiving, via the UI, one or more selections of respective one or more stylistic features from the list of stylistic features. (“The method 600 includes receiving data at 604 uploaded by the user and storing the data on the image server for cloud-based access. The method 600 includes providing at 606 a user interface to the user that enables the user to organize and manage the data collection. The method 600 includes receiving at 608 a thumbnail selection approval from the user, wherein the thumbnail comprises a screen capture from one or more data objects within the data collection. The method 600 includes generating at 610 a unique code that provides instructions to a personal device to access the data collection. The method 600 includes merging at 612 the unique code and the thumbnail selection to generate a scannable thumbnail. The data uploaded by the user may include one or more independent data objects such as videos, images, written works, numerical data, historical data, and so forth. The one or more data objects may be grouped together to generate a data collection.” [0063-0064] “FIGS. 8A-8F illustrate screenshots of an example user interface 800. FIG. 8G is an example scannable thumbnail as described herein. The user interface 800 supports user interactions with the media platform 102 supported by the media server 104. The user interface 800 enables users to upload media to the media library 112, edit media stored on the media library 112, select a thumbnail, cause a unique code to be generated, select media to be associated with the unique code, generate a scannable thumbnail, and so forth as discussed herein.” [0068]) 16. Claim 15 is similar in scope to claim 5, and thus is rejected under similar rationale. 17. Claim 17 is similar in scope to claim 1, and thus is rejected under similar rationale. Willis additionally teaches A non-transitory computer readable storage medium comprising instructions for a server that, when executed by a processing device, cause the processing device to perform operations (“Computing device 900 includes one or more processor(s) 902, one or more memory device(s) 904, one or more interface(s) 906, one or more mass storage device(s) 908, one or more Input/output (I/O) device(s) 910, and a display device 930 all of which are coupled to a bus 912. Processor(s) 902 include one or more processors or controllers that execute instructions stored in memory device(s) 904 and/or mass storage device(s) 908. Processor(s) 902 may also include various types of computer-readable media, such as cache memory. Memory device(s) 904 include various computer-readable media, such as volatile memory (e.g., random access memory (RAM) 914) and/or nonvolatile memory (e.g., read-only memory (ROM) 916). Memory device(s) 904 may also include rewritable ROM, such as Flash memory.” [0078-0079] “Non-transitory computer readable storage media for storing instructions to be executed by one or more processors,” claim 16) 18. Claim 18 is similar in scope to claim 2, and thus is rejected under similar rationale. 19. Claim 21 is similar in scope to claim 5, and thus is rejected under similar rationale. 20. Claim 23 is similar in scope to claim 8, and thus is rejected under similar rationale. Willis additionally teaches A system comprising: a memory device; and a processing device coupled to the memory device, the processing device to perform operations (“Computing device 900 includes one or more processor(s) 902, one or more memory device(s) 904, one or more interface(s) 906, one or more mass storage device(s) 908, one or more Input/output (I/O) device(s) 910, and a display device 930 all of which are coupled to a bus 912. Processor(s) 902 include one or more processors or controllers that execute instructions stored in memory device(s) 904 and/or mass storage device(s) 908. Processor(s) 902 may also include various types of computer-readable media, such as cache memory.” [0078] “A system comprising one or more processors for executing instructions stored on non-transitory computer readable storage media,” claim 11) 21. Claim 24 is similar in scope to claim 9, and thus is rejected under similar rationale. 22. Claim(s) 3, 4, 6, 10-14, 16, 19, 20, 22 and 25-28 is/are rejected under 35 U.S.C. 103 as being unpatentable over Willis et al. (US 2023/0104396) and Cragg et al. (US 2024/0273670 A1), as applied to claims 1, 8, 9, 17, 18, 23 and 24 above, and further in view of O’Neill (US 2024/0412542 A1). 23. With reference to claim 3, the combination of Willis and Cragg does not explicitly teach causing the AI generative model to process the textual prompt is performed responsive to determining that the textual prompt satisfies a content appropriateness condition. This is what O’Neill teaches (“the processor may perform multi-modal analysis operations to determine the attributes in the audio component of a selected media content segment. For example, in cases where audio is part of multi-modal content (e.g., audio-video), the processor may integrate audio analysis with visual and textual analysis for a more comprehensive understanding. In some embodiments, the processor may use multi-modal LXMs to correlate and synthesize insights from different content types.” [0261] “the processor may be configured to load and evaluate rules, select rules, search the ToI knowledge repository, determine whether a condition of the selected rule has been satisfied, check the rule action time period, execute corresponding rule actions based on the properties of the media content and publisher perform rule actions, engage with media content, perform promotional activities, promote social media content, create a promotional order, and/or manage ads based on content publisher behavior.” [0562] “the processor may retrieve the relevant requirements relating to the media content (e.g., from an external source, from the product knowledge repository). In block 3212, the processor may determine whether the media content is published in accordance with the relevant requirements. This determination may be based on a threshold (e.g., up to three associations with competitors may be allowed, but no inappropriate topics are allowed).” [0589]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of O’Neill into the combination of Willis and Cragg, in order to perform a responsive action. 24. With reference to claim 4, the combination of Willis and Cragg does not explicitly teach determining whether the textual prompt satisfies the content appropriateness condition further comprises: comparing the textual prompt to an allowlist of prompt terms. This is what O’Neill teaches (“the processor may perform multi-modal analysis operations to determine the attributes in the audio component of a selected media content segment. For example, in cases where audio is part of multi-modal content (e.g., audio-video), the processor may integrate audio analysis with visual and textual analysis for a more comprehensive understanding. In some embodiments, the processor may use multi-modal LXMs to correlate and synthesize insights from different content types.” [0261] “the processor may be configured to load and evaluate rules, select rules, search the ToI knowledge repository, determine whether a condition of the selected rule has been satisfied, check the rule action time period, execute corresponding rule actions based on the properties of the media content and publisher perform rule actions, engage with media content, perform promotional activities, promote social media content, create a promotional order, and/or manage ads based on content publisher behavior.” [0562] “the processor may retrieve the relevant requirements relating to the media content (e.g., from an external source, from the product knowledge repository). In block 3212, the processor may determine whether the media content is published in accordance with the relevant requirements. This determination may be based on a threshold (e.g., up to three associations with competitors may be allowed, but no inappropriate topics are allowed).” [0589] “the processor may use image recognition techniques (pattern recognition, geometric analysis, etc.) to compare the observed shape of the product against a predefined list of shapes and/or categorize the basic form of the product into a predefined shape.” [0595]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of O’Neill into the combination of Willis and Cragg, in order to perform a responsive action. 25. With reference to claim 6, Willis teaches each of the one or more thumbnail images (“The method 600 includes receiving at 608 a thumbnail selection approval from the user, wherein the thumbnail comprises a screen capture from one or more data objects within the data collection. The method 600 includes generating at 610 a unique code that provides instructions to a personal device to access the data collection. The method 600 includes merging at 612 the unique code and the thumbnail selection to generate a scannable thumbnail.” [0063]) The combination of Willis and Cragg does not explicitly teach providing input to a second trained AI model; and obtaining one or more outputs of the second trained AI model, the one or more outputs of the second AI model indicating a probability comprising an inappropriate content. This is what O’Neill teaches (“The term “deep neural network” may be used herein to refer to a neural network that implements a layered architecture in which the output/activation of a first layer of nodes becomes an input to a second layer of nodes, the output/activation of a second layer of nodes becomes an input to a third layer of nodes, and so on.” [0068] “The term “sequence data processing” may be used herein to refer to techniques or technologies for handling ordered sets of tokens in a manner that preserves their original sequential relationships and captures dependencies between various elements within the sequence. The resulting output may be a probabilistic distribution or a set of probability values, each corresponding to a “possible succeeding token” in the existing sequence.” [0082] the processor may suggest values to populate the fields for brand names or product names 2702, topics that are of interest to the user 2708, territories and/or regions that are of interest to the user 2710, list of competitors 2712, list of inappropriate topics with which the user does not want to be associated 2714, and list of values which are important to the user 2716 (e.g., as part of block 2622 with reference to FIG. 26).” [0554]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of O’Neill into the combination of Willis and Cragg, in order to perform a responsive action. 26. With reference to claim 10, the combination of Willis and Cragg does not explicitly teach the list of stylistic features is personalized based on a user profile associated with the user. This is what O’Neill teaches (“access user profiles to gather context (e.g., interests, typical attire, etc.), use this contextual information to further refine search accuracy, automatically scrape and process additional relevant information from external sources, update the database with the new data (e.g., to improve future searches, etc.), and/or generate a report summarizing the identified products/brands, their context, confidence scores, etc. In some embodiments, the components may be configured to regularly update machine learning models, user profiles, product databases, etc. based on recent trends and results of the content analysis.” [0131] “Further examples of responsive actions include the processor launching advertising campaigns tailored to specific audience segments identified through media content analysis, updating or adjusting a content strategy based on the analysis of media content segment, using the determined attributes from the media content segments to personalize content for users on a platform,” [0201] “the processor may be configured to recognize specific topics of interest to the user (e.g., industry-specific terminology, branded content, personalized themes, etc.) via custom keyword lists or machine learning models trained on specialized content.” [0242]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of O’Neill into the combination of Willis and Cragg, in order to perform a responsive action. 27. With reference to claim 11, the combination of Willis and Cragg does not explicitly teach the list of stylistic features is translated to a natural language specified by a user profile associated with the user. This is what O’Neill teaches (“the components may be configured to load machine learning models for content analysis and natural language processing (NLP), initialize databases for storing user profiles and product/brand information, … access user profiles to gather context (e.g., interests, typical attire, etc.), use this contextual information to further refine search accuracy, automatically scrape and process additional relevant information from external sources, update the database with the new data (e.g., to improve future searches, etc.), and/or generate a report summarizing the identified products/brands, their context, confidence scores, etc. In some embodiments, the components may be configured to regularly update machine learning models, user profiles, product databases, etc. based on recent trends and results of the content analysis.” [0131] “Further examples of responsive actions include the processor launching advertising campaigns tailored to specific audience segments identified through media content analysis, updating or adjusting a content strategy based on the analysis of media content segment, using the determined attributes from the media content segments to personalize content for users on a platform,” [0201] “the processor may be configured to recognize specific topics of interest to the user (e.g., industry-specific terminology, branded content, personalized themes, etc.) via custom keyword lists or machine learning models trained on specialized content.” [0242]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of O’Neill into the combination of Willis and Cragg, in order to perform a responsive action. 28. With reference to claim 12, the combination of Willis and Cragg does not explicitly teach receiving, via the UI, a list refresh command; and responsive to receiving the list refresh command, refreshing the list of stylistic features. This is what O’Neill teaches (“analyze each product image to determine whether it contains additional relevant or irrelevant elements (e.g., product alone vs. product with a model), store only relevant images or extract the essential parts of the image for the database, and periodically re-scrape websites to capture any new or updated products and logo and regularly update the database with new categories and products, query the database in response to detecting a new product image to determine whether the database includes similar images or logos, use the stored information for quick logo recognition in response to determining that the database includes similar images or logos, and perform a web search for logo recognition in response to determining that the database does not include similar images or logos.” [0117] “access user profiles to gather context (e.g., interests, typical attire, etc.), use this contextual information to further refine search accuracy, automatically scrape and process additional relevant information from external sources, update the database with the new data (e.g., to improve future searches, etc.), and/or generate a report summarizing the identified products/brands, their context, confidence scores, etc. In some embodiments, the components may be configured to regularly update machine learning models, user profiles, product databases, etc. based on recent trends and results of the content analysis.” [0131] “access reliable and current external sources for trend monitoring, extract audio and textual content from one or more video streams or files, analyze the extracted content to identify primary themes and keywords, cross-reference these themes with a pre-existing database of contextually relevant terms, scan external sources (e.g., news, social media, etc.) for emerging trends and topics, update the database with new trends, topics and associations, integrate the newly identified trends and topics into the analysis models, update, refine, or fine tune the models to recognize and correctly interpret the new terms within the video content, apply the updated models to analyze the content for each new video, apply the updated models to analyze the content, identify and categorize themes and subjects based on current global contexts, classify video content based on evolving definitions and understandings of key terms (e.g., “armed conflict,” etc.), and adjust the classifications as global contexts and terminology evolve or change.” [0135] “a human operator may initiate the request to add media content via user input (e.g., clicking through options in a graphical user interface (GUI), typing commands into a command line, using voice commands in a voice recognition system, etc.).” [0175] “the processor may be configured to recognize specific topics of interest to the user (e.g., industry-specific terminology, branded content, personalized themes, etc.) via custom keyword lists or machine learning models trained on specialized content.” [0242] “the processor may receive the query through a user interface that allows for users to input specific search terms, through an API from another computing device, through a remote procedure call from a component configured to track certain topics or trends, etc. The query may, for example, seek information about the latest discussions and sentiment surrounding new skincare trends (e.g., skin flooding, natural oils, under eye patches, face self-tanners, etc.).” [0297]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of O’Neill into the combination of Willis and Cragg, in order to perform a responsive action. 29. Claims 13 and 14 are similar in scope to claims 3 and 4, and they are rejected under similar rationale. 30. Claim 16 is similar in scope to claim 6, and thus is rejected under similar rationale. 31. Claims 19 and 20 are similar in scope to claims 3 and 4, and they are rejected under similar rationale. 32. Claim 22 is similar in scope to claim 6, and thus is rejected under similar rationale. 33. Claims 25-28 are similar in scope to claims 10-13, and they are rejected under similar rationale. Conclusion 34. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michelle Chin whose telephone number is (571)270-3697. The examiner can normally be reached on Monday-Friday 8:00 AM-4:30 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http:/Awww.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Kent Chang can be reached on (571)272-7667. The fax phone number for the organization where this application or proceeding is assigned is (571)273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https:/Awww.uspto.gov/patents/apply/patent- center for more information about Patent Center and https:/Awww.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHELLE CHIN/ Primary Examiner, Art Unit 2614
Read full office action

Prosecution Timeline

Show 4 earlier events
Nov 25, 2025
Applicant Interview (Telephonic)
Dec 10, 2025
Response Filed
Mar 02, 2026
Final Rejection mailed — §103
May 21, 2026
Examiner Interview Summary
May 21, 2026
Applicant Interview (Telephonic)
Jun 01, 2026
Request for Continued Examination
Jun 04, 2026
Response after Non-Final Action
Jun 11, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12702485
FEEDBACK FOR SURGICAL ROBOTIC SYSTEM WITH VIRTUAL REALITY
2y 5m to grant Granted Aug 11, 2026
Patent 12694623
3D PHOTOS
3y 8m to grant Granted Jul 28, 2026
Patent 12694618
3D OBJECT IMAGING METHOD AND 3D OBJECT IMAGING SYSTEM
2y 7m to grant Granted Jul 28, 2026
Patent 12682540
GRAPHICS PROCESSING
2y 6m to grant Granted Jul 14, 2026
Patent 12675917
IDENTITY-PRESERVING IMAGE GENERATION USING DIFFUSION MODELS
3y 1m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
86%
Grant Probability
97%
With Interview (+11.2%)
2y 2m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 650 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month