DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Remarks
This action is in response to the applicant’s response filed 30 June 2026, which is in response to the USPTO office action mailed 3 April 2026. Claims 1 and 20 are amended. Claim 2 is cancelled. Claim 21 is added. Claims 1 and 3-21 are currently pending.
Response to Arguments
With respect to the 35 USC §103 rejection of claims 1 and 3-20, the applicant’s arguments are moot in view of a new grounds of rejection, as necessitated by the applicant's amendments.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 and 3-20 are rejected under 35 U.S.C. 103 as being unpatentable over Atkins et al., US 2012/0272126 A1 (hereinafter “Atkins”) in view of Barrett et al., US 2017/0230312 A1 (hereinafter “Barrett”) in further view of TSENG, US 2012/0076367 A1 (hereinafter “Tseng”) in further view of Kottur et al., US 2023/0401170 A1 (hereinafter “Kottur”) in further view of XIE et al., US 2023/0394855 A1 (hereinafter “Xie”) in further view of Raya et al., US 12,039,653 B1 (hereinafter “Raya”).
Claim 1: Atkins teaches a method of displaying a digital photo collection, the method comprising:
automatically extracting content features from photos of the digital photo collection (Atkins, [0018] note electronically access a source such as database of photos and access a collection of photos, [Fig. 3], [0056] note photo module 260 is configured to track metadata obtained via electronically analyzing or observing the subject(s) of each photo);
creating metadata tags for each photo as a function of corresponding extracted content features for the photo (Atkins, [Fig. 4], [0064] note theme parameter 304 tracks a theme identified by the author or automatically generated via analysis of the content metadata of the photos);
storing tagged photos in a database (Atkins, [0018] note electronically access a source such as database of photos and access a collection of photos, [0036] note labeling information, [0053] note metadata manager 225 is configured as a database to comprehensively store and manage metadata produced via operation of compilation manager 100);
receiving a request; automatically determining a search parameter from the request (Atkins, [0036] note the search function 150 enables an author to search among photos or other media elements within a general source 146 (part of content selector 130) or within an already selected group of photos or other media elements. The search is performed via keywords or other searching protocols known in the art);
automatically connecting to the database over a network (Atkins, [0088] note computer system 600 comprises a first computer 602, a content service provider 604, a compilation service provider 606, an output provider 608, and a network communication link 610);
automatically searching the database for metadata tags and digital photos matching the search parameter (Atkins, [0036] note The search is performed via keywords or other searching protocols known in the art); and
automatically displaying on the digital display one or more matching photos obtained from the database over the network connection (Atkins, [0051] note the compilation manager 100 also includes an output monitor 200, as shown in FIG. 2. In general terms, output monitor 200 enables selection of the type of output of the first media compilation 134, the second media compilation 188, or successive derivations 190. This output is generally complementary to the selected format of the respective media compilations. In one embodiment, output monitor 200 comprises a photobook function 210, [Fig. 7], [0088] note output provider 608, and a network communication link 610) and
automatically generating a unique narrative matching the request and the one or more matching photos displayed on the digital display, wherein the unique narrative comprises a unique storyline guided by displayed image content, photo metadata, and settings provided by the viewer (Atkins, [0015] note building a media compilation, [0022] note composing and editing includes selecting a format, such as a photobook, slideshow, collage and arranging the photos within that selected format. This process includes several aspects, such as, but not limited to, choosing: (1) how many photos will appear on a single page: (2) the relative sizes of the photos; (3) their orientation; (4) a sequence of the photos; and/or (5) how the photos are grouped together. In one aspect, the author can choose a predetermined format according to one or more themes, such as a birthday, sports season, wedding, etc… In some embodiments, an automated process can be applied to automatically populate the fields in the predetermined format with photos that are automatically selected according to their content metadata).
Atkins does not explicitly teach on a digital picture frame including a digital display mounted within a frame, a microphone and speaker connected to the frame, and a network connection module; confirming the created metadata tags for the photo using photo metadata of the photo or a related photo; assigning a confidence level to the created metadata tags as a function of the confirming step; determining additional analysis is needed for the metadata tags of one of the photos; automatically initiating a verbal interaction through the speaker with a viewer of the one of the photos to request and receive a content detail of the one of the photos; improving the metadata tags of the one of the photos as a function of the verbal interaction and received content detail; increasing an assigned confidence level for the metadata tags of the one of the photos as a function of the verbal interaction and the received content detail; from a viewer of the digital picture frame; spoken through the speaker, and at least partially fictitious.
However, Barrett teaches on a digital picture frame including a digital display mounted within a frame, a microphone and speaker connected to the frame, and a network connection module; and from a viewer of the digital picture frame (Barrett, [0012] note one or more embodiments described herein may be implemented, in whole or in part, on… digital picture frames, [0014] note an interface that can simulate a chat conversation with a human, [0022] note chatbot 102 can handle a chat input 11 by providing a chat response 12, which can include at least one of (1) a question for the user to answer, and/or (2) information determined in response to the chat input 11, [0023] note Based on context and other information, the chatbot 102 can make an initial determination of whether a suitable response (e.g., information or follow on question) is determinable given information and resources available to the chatbot 102, [0042] note the user interface 110 can include programmatic voice interaction, where vocal human input can be recognized as chat input 11, and chat responses 12 can be output as voice synthesis. The interaction between user and chatbot 102 can be conducted through, for example, a telephone, microphone/speaker).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the media compilations of Atkins with the chat engine implemented on a digital picture frame of Barrett according to known methods (i.e. implementing a chat engine which can simulate a chat conversation with a human on a digital picture frame). Motivation for doing so is that the chatbot can provide information, identify information sources, and/or ask questions for the user to answer to progressively advance the query from the user (Barrett, [0008]).
Atkins and Barrett do not explicitly teach confirming the created metadata tags for the photo using photo metadata of the photo or a related photo; assigning a confidence level to the created metadata tags as a function of the confirming step; determining additional analysis is needed for the metadata tags of one of the photos; automatically initiating a verbal interaction through the speaker with a viewer of the one of the photos to request and receive a content detail of the one of the photos; improving the metadata tags of the one of the photos as a function of the verbal interaction and received content detail; increasing an assigned confidence level for the metadata tags of the one of the photos as a function of the verbal interaction and the received content detail; spoken through the speaker, and at least partially fictitious.
However, Tseng teaches confirming the created metadata tags for the photo using photo metadata of the photo or a related photo; assigning a confidence level to the created metadata tags as a function of the confirming step (Tseng, [0018] note automated media tagging capabilities, [0021] note a facial recognition or matching algorithm that returns a matching score and comparing the score to a threshold value, [0022] note the auto-tagging process may apply spatio-temporal matching (206) to adjust the matching scores for the potential matches (207)…. the auto-tagging process may determine location data of an image file by accessing geographic location data (e.g. GPS coordinates), user check-in activity data, and/or event data associated with the image file, or based on one or more previous matches already identified in the image file, where the matched user's location is known with some threshold degree of confidence).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the media compilations of Atkins and Barrett with the auto-tagging process based on a matching score of Tseng according to known methods (i.e. auto-tagging media based on comparing a matching score to a threshold value). Motivation for doing so is that this improves the accuracy of the tagging process (Tseng, [0003]).
Atkins, Barrett and Tseng do not explicitly teach determining additional analysis is needed for the metadata tags of one of the photos; automatically initiating a verbal interaction through the speaker with a viewer of the one of the photos to request and receive a content detail of the one of the photos; improving the metadata tags of the one of the photos as a function of the verbal interaction and received content detail; increasing an assigned confidence level for the metadata tags of the one of the photos as a function of the verbal interaction and the received content detail; spoken through the speaker, and at least partially fictitious.
However, Kottur teaches determining additional analysis is needed for the metadata tags of one of the photos; automatically initiating a verbal interaction through the speaker with a viewer of the one of the photos to request and receive a content detail of the one of the photos; improving the metadata tags of the one of the photos as a function of the verbal interaction and received content detail; increasing an assigned confidence level for the metadata tags of the one of the photos as a function of the verbal interaction and the received content detail (Kottur, [0145] note the user may request to add metadata to an album associated with the user's digital memories. As an example and not by way of limitation, the user may say “hey assistant, show me photos of my ‘Summer 2019’ album.” The assistant system 140 may reply “I don't see an album called ‘Summer 2019’, but here are some albums with photos from Summer 2019.” The assistant system 140 may additionally show two albums with photos taken in the summer months of 2019: “beach 2019” and “Las Vegas sun”. The user may then say “for the album titled ‘beach 2019’, add the description, ‘photos from Long Beach Island in the summer of 2019.’” The assistant system 140 may then reply “okay, I'll add the description, ‘photos from Long Beach Island in the summer of 2019’ to the album.”, [0086] note dialog manager 216 may implement reinforcement learning frameworks to improve the dialog optimization, [0115] note a conversational understanding reinforcement engine (CURE) tracker 420. In particular embodiments, the CURE tracker 420 may be a personalized learning process to improve the determination of the state candidates by the dialog state tracker 218 under different contexts using real-time user feedback).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the tagging of Atkins, Barrett and Tseng with the user request to add metadata of Kottur according to known methods (i.e. adding metadata to photos based on a user request). Motivation for doing so is that this provides a technical advantage which includes increasing the degree of users engaging with the assistant system by accurately retrieving their priming digital memories and proactively providing users with the related digital memories responsive to their queries, thereby generating multi-turn conversations (Kottur, [0011]).
Atkins, Barrett, Tseng and Kottur do not explicitly teach spoken through the speaker, and at least partially fictitious.
However, Xie teaches at least partially fictitious (Xie, [Fig. 1], [Fig. 2], [Fig. 3] note 312, [0018] note A rich semantic representation of an input image, such as image tags, object attributes and locations, and captions, is constructed, [0021] note A generative language model 140 generates a plurality of image story caption candidates 144, which includes story captions 141, 142, and 143, from visual information 120, which includes, from visual clues 130, [0027] note story caption 154 is then paired with image 102 in paired set 160, [0028] note architecture 100 generates story caption 312).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the media compilations of Atkins, Barrett, Tseng and Kotur with the story captions of Xie according to known methods (i.e. generating story captions for input images). Motivation for doing so is that, relative to captions generated by existing systems, “story” captions or paragraph captions generated by the present disclosure are longer, more descriptive, more accurate, and include additional details about the input image or additional information related to the input image. In this manner, the input image is described more thoroughly and precisely (Xie, [0016]).
Atkins, Barrett, Tseng, Kottur and Xie do not explicitly teach spoken through the speaker.
However, Raya teaches this (Raya, [Fig. 3], [Col. 6 Lines 13-25] note input data can be provided to a narration model 302, which can use the input data to generate narrative text… The narrative text generated by the narration model 302 can also be provided to a text-to-speech model 308, which it can use to generate narrative speech, [Fig. 4], [Col. 8 Lines 50-51] note the example narrative text discussed above and depicted as narrative text 402).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the generated image story captions of Atkins, Barrett, Tseng, Kottur and Xie with the narrative speech of Raya according to known methods (i.e. generating narrative speech based on image story captions). Motivation for doing so is that this provides for easy and efficient content generation (Raya, [Col. 2 Lines 23-24]).
Claim 3: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, further comprising:
automatically providing a second verbal interaction with the viewer of the digital picture frame upon receiving the request, wherein the second verbal interaction comprises a plurality of automated back-and-forth conversational iterations between the frame speaker and the viewer of the digital picture frame to establish the search parameter, the plurality of automated back-and-forth conversational iterations comprising more than one spoken statement or spoken question from the frame speaker (Kottur, [0145] note the user may request to add metadata to an album associated with the user's digital memories. As an example and not by way of limitation, the user may say “hey assistant, show me photos of my ‘Summer 2019’ album.” The assistant system 140 may reply “I don't see an album called ‘Summer 2019’, but here are some albums with photos from Summer 2019.” The assistant system 140 may additionally show two albums with photos taken in the summer months of 2019: “beach 2019” and “Las Vegas sun”. The user may then say “for the album titled ‘beach 2019’, add the description, ‘photos from Long Beach Island in the summer of 2019.’” The assistant system 140 may then reply “okay, I'll add the description, ‘photos from Long Beach Island in the summer of 2019’ to the album.”).
Claim 4: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, wherein the confirming the created metadata tags for the photo comprises analyzing a location, date, time, and occasion setting classification of the photo (Tseng, [0020] note a photo shot by a digital camera may contain metadata relating to file size, resolution, time stamp, name of the camera maker, and/or location (e.g., GPS) coordinates, [0022] note the auto-tagging process may determine location data of an image file by accessing geographic location data (e.g. GPS coordinates), user check-in activity data, and/or event data associated with the image file, or based on one or more previous matches already identified in the image file, where the matched user's location is known with some threshold degree of confidence).
Claim 5: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, wherein the confirming the created metadata tags for the photo comprises automatically identifying location, activity, and community member involvement for the photo (Tseng, [0021] note the auto-tagging process may identify one or more faces in the image file that correspond to other users of social networking system or individuals generally, [0022] note the auto-tagging process may determine location data of an image file by accessing geographic location data (e.g. GPS coordinates), user check-in activity data, and/or event data associated with the image file).
Claim 6: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, wherein the confirming the created metadata tags for the photo comprises mining correlations within metadata tags by determining relationships between different metadata tags of a plurality of the tagged photos (Tseng, [0022] note the auto-tagging process can determine that a photo is associated with location Golden Gate Bridge and time stamp Oct. 8, 2009 by an event tagged by a user to the photo (e.g., "Fleet Week 2009, Golden Gate Bridge, Oct. 8, 2009"), or a title of a photo album that the photo belongs to (e.g., "Golden Gate Bridge, Oct. 8, 2009"). For example, if a user is already identified in an image file (e.g., having a correlation coefficient=1.0), the image file has a time stamp of Oct. 8, 2009, and the user has an event "Fleet Week 2009, Golden Gate Bridge, Oct. 8, 2009", then the auto-tagging process can determine that the image file is associated with an event "Fleet Week 2009" and with a location "Golden Gate Bridge.”).
Claim 7: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, wherein the confirming the created metadata tags for the photo comprises determining a relationship between different metadata tags of the photo or the related photo (Tseng, [0022] note the auto-tagging process can determine that a photo is associated with location Golden Gate Bridge and time stamp Oct. 8, 2009 by an event tagged by a user to the photo (e.g., "Fleet Week 2009, Golden Gate Bridge, Oct. 8, 2009"), or a title of a photo album that the photo belongs to (e.g., "Golden Gate Bridge, Oct. 8, 2009"). For example, if a user is already identified in an image file (e.g., having a correlation coefficient=1.0), the image file has a time stamp of Oct. 8, 2009, and the user has an event "Fleet Week 2009, Golden Gate Bridge, Oct. 8, 2009", then the auto-tagging process can determine that the image file is associated with an event "Fleet Week 2009" and with a location "Golden Gate Bridge.”).
Claim 8: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, wherein the confirming the created metadata tags for the photo comprises mining correlations between different metadata tags of the photo or the related photo (Tseng, [0023] note if a potential match has a correlation coefficient of 0.75 as calculated by the facial recognition computer software, and the potential match in close spatio-temporal proximity to the location data associated with the image file, the auto-tagging process can increase the match score).
Claim 9: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, wherein the confirming the created metadata tags for the photo comprises comparing learned context data stored for a date of the photo or the related photo (Tseng, [0023] note a user corresponding to a potential match in the selected subset of potential matches may have checked in at the Golden Gate Bridge using a geo-social networking client on Oct. 8, 2009. In connection with check-ins using a geo-social networking client application, the location database may store GPS location (e.g., 37.degree. 49'09.15'' N, 122.degree. 28'45 11'' W) on Oct. 8, 2009, as provided by the client application. Alternatively or additionally, the user may have configured for a calendar entry with a location of "Golden Gate Bridge" on the day Oct. 8, 2009." Still further, the user may have configured, or registered as attending, an event having a location near or at the Golden Gate Bridge).
Claim 10: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, wherein the confirming the created metadata tags for the photo comprises analyzing tracked location information of a mobile device used to take the photo (Tseng, [0023] note In connection with check-ins using a geo-social networking client application, the location database may store GPS location (e.g., 37.degree. 49'09.15'' N, 122.degree. 28'45 11'' W) on Oct. 8, 2009, as provided by the client application).
Claim 11: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 10, further comprising automatically learning a location, activity, and community member involvement on a per photo basis from the tracked location information of the mobile device (Tseng, [0022] note the auto-tagging process can determine that a photo is associated with location Golden Gate Bridge and time stamp Oct. 8, 2009 by an event tagged by a user to the photo (e.g., "Fleet Week 2009, Golden Gate Bridge, Oct. 8, 2009"), or a title of a photo album that the photo belongs to (e.g., "Golden Gate Bridge, Oct. 8, 2009"). For example, if a user is already identified in an image file (e.g., having a correlation coefficient=1.0), the image file has a time stamp of Oct. 8, 2009, and the user has an event "Fleet Week 2009, Golden Gate Bridge, Oct. 8, 2009", then the auto-tagging process can determine that the image file is associated with an event "Fleet Week 2009" and with a location "Golden Gate Bridge.”).
Claim 12: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, wherein the confirming the created metadata tags for the photo comprises computing a likelihood of a given person being at a given location or at a given time using the photo metadata (Tseng, [0022] note the auto-tagging process can determine that a photo is associated with location Golden Gate Bridge and time stamp Oct. 8, 2009 by an event tagged by a user to the photo (e.g., "Fleet Week 2009, Golden Gate Bridge, Oct. 8, 2009"), or a title of a photo album that the photo belongs to (e.g., "Golden Gate Bridge, Oct. 8, 2009"). For example, if a user is already identified in an image file (e.g., having a correlation coefficient=1.0), the image file has a time stamp of Oct. 8, 2009, and the user has an event "Fleet Week 2009, Golden Gate Bridge, Oct. 8, 2009", then the auto-tagging process can determine that the image file is associated with an event "Fleet Week 2009" and with a location "Golden Gate Bridge.”).
Claim 13: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, wherein the confirming the created metadata tags for the photo comprises reducing photo content ambiguities as a function of photo location metadata and photo date and time metadata (Tseng, [0020] note a photo shot by a digital camera may contain metadata relating to file size, resolution, time stamp, name of the camera maker, and/or location (e.g., GPS) coordinates, [0023] note if a potential match has a correlation coefficient of 0.75 as calculated by the facial recognition computer software, and the potential match in close spatio-temporal proximity to the location data associated with the image file, the auto-tagging process can increase the match score).
Claim 14: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, wherein the confirming the created metadata tags for the photo comprises clustering photos using common detected photo content, photo location metadata, and photo date and time metadata (Atkins, [0064] note clustering parameter 302 tracks which photos or sets of photos are clustered together within a media compilation while the theme parameter 304 tracks a theme).
Claim 15: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, wherein the confirming the created metadata tags for the photo comprises analyzing an activity, a location, a person, a season, or an outfit worn in the photo (Tseng, [0021] note the auto-tagging process may identify one or more faces in the image file that correspond to other users of social networking system or individuals generally, [0022] note the auto-tagging process may determine location data of an image file by accessing geographic location data (e.g. GPS coordinates), user check-in activity data, and/or event data associated with the image file).
Claim 16: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, wherein the confirming the created metadata tags for the photo comprises analyzing and correlating more than one extracted photo content (Atkins, [0021] note the auto-tagging process may identify one or more faces in the image file that correspond to other users of social networking system or individuals generally (202)… facial recognition computer software can calculate a correlation coefficient between a potential match and an identified face, where the correlation coefficient ranges from 0.0 ("no correlation at all") to 1.0 ("perfect match")).
Claim 17: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, wherein the confirming the created metadata tags for the photo comprises comparing a first extracted content feature from the photo to a second extracted content feature from the photo (Atkins, [0021] note the auto-tagging process may identify one or more faces in the image file that correspond to other users of social networking system or individuals generally (202)… facial recognition computer software can calculate a correlation coefficient between a potential match and an identified face, where the correlation coefficient ranges from 0.0 ("no correlation at all") to 1.0 ("perfect match")).
Claim 18: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim1, wherein extracted content features a location, a season, a weather content, a time of day, a date, an activity content, and one or more attributes of a person of the photo (Atkins, [0019] note ach photo includes a metadata tag storing this information, which may include a time or date the photo was taken, a location (e.g. GPS) the photo was taken, [0022] note themes, such as a birthday, sports season, wedding, etc, [0056] note observing the subject(s) of each photo. The photo module 260 comprises a facial parameter 272, a gender parameter 274).
Claim 19: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, wherein the verbal interaction comprises at least one automated back-and-forth conversational iteration with the viewer of the one of the photos (Kottur, [0145] note the user may request to add metadata to an album associated with the user's digital memories. As an example and not by way of limitation, the user may say “hey assistant, show me photos of my ‘Summer 2019’ album.” The assistant system 140 may reply “I don't see an album called ‘Summer 2019’, but here are some albums with photos from Summer 2019.” The assistant system 140 may additionally show two albums with photos taken in the summer months of 2019: “beach 2019” and “Las Vegas sun”. The user may then say “for the album titled ‘beach 2019’, add the description, ‘photos from Long Beach Island in the summer of 2019.’” The assistant system 140 may then reply “okay, I'll add the description, ‘photos from Long Beach Island in the summer of 2019’ to the album.”).
Claim 20: Atkins, Barrett, Tseng, Kottur, Xie and Raya teach the method of Claim 1, wherein the at least partially fictitious and unique storyline includes a character, a setting, and a plot, each based at least partially on or corresponding at least partially to the displayed image content (Xie, [Fig. 1], [Fig. 2], [Fig. 3] note 312, [0018] note A rich semantic representation of an input image, such as image tags, object attributes and locations, and captions, is constructed, [0021] note A generative language model 140 generates a plurality of image story caption candidates 144, which includes story captions 141, 142, and 143, from visual information 120, which includes, from visual clues 130, [0027] note story caption 154 is then paired with image 102 in paired set 160, [0028] note architecture 100 generates story caption 312).
Claim 21 is rejected under 35 U.S.C. 103 as being unpatentable over Atkins, Barrett, Tseng, Kottur, Xie and Raya in further view of Benedetto et al., US 2024/0299851 A1 (hereinafter “Benedetto”).
Claim 21: Atkins, Barrett, Tseng, Kottur, Xie and Raya do not explicitly teach the method of Claim 20, further comprising automatically adjusting the partially fictitious and unique spoken storyline of the unique narrative in response to responsive vocal utterances or voice prompts from the viewer viewing the one or more matching photos displayed on the digital display.
However, Benedetto teaches this (Benedetto, [0020] note constraints are defined and applied to the AI-based storyboard generation process to assist in steering a direction of the story conveyed by the storyboard and/or to avoid inclusion of unwanted types of content within the storyboard, [0027] note the AI-based storyboard generation process can be used to automatically generate a short book, a novel, a movie, a game, or essentially any other story-conveying product based on seed input information and user-specified constraints. In some embodiments, the AI-based storyboard generation process is moderated by constraints and/or prompts provided by a single user so that the storyboard is generated and steered by a single person, [0058] note the user-supplied steering input received in the operation 621 is in an audio format).
It would have been obvious to one of ordinary skill in the art at the effective filing date of the application to combine the generated image story captions of Atkins, Barrett, Tseng, Kottur, Xie and Raya with the user-supplied steering input of Benedetto according to known methods (i.e. receiving user-supplied steering prompts for generating image story captions). Motivation for doing so is that this assists in steering a direction of the story conveyed to avoid inclusion of unwanted types of content (Benedetto, [0020]).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Giuseppi Giuliani whose telephone number is (571)270-7128. The examiner can normally be reached Monday-Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kavita Stanley can be reached at (571)272-8352. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/GIUSEPPI GIULIANI/Primary Examiner, Art Unit 2153