Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
Claims 1-3, 5, 6, 8, 13, 15-20 are amended, claim 14 is canceled, claim 21 is added. Claims 1-13, and 15-21 are presented for examination.
Response to Arguments
Applicant’s arguments filed on 06/02/2024 have been reviewed. Following are the responses to the amendments.
Rejection under 35 U.S.C. 101
Applicant arguments have been considered and are persuasive, hence the rejection under 35 U.S.C. 101 is withdrawn.
Rejection under 35 U.S.C. 103
“Qiu and Huang, taken alone or in hypothetical combination, fail to teach or suggest all of the elements of independent claims 1, 8, and 16.”
Applicant argues that “Baeuml and Huang, taken alone or in hypothetical combination, fail to teach or suggest all of the elements of independent claims 1, 8, and 16” including the limitations direction to receiving images of a guest during a visit to an amusement park. The examiner agrees that Qiu and Huang do not expressly teach these amended limitations. Therefore, the prior rejection has been withdrawn. However, upon further consideration, a new ground of rejection has been made over Qiu, Cronin and Kachuee for claim 1, and Qui, Cronin, Kachuee, and Huang for claims 8 and 16.
“Dependent Claims 6-7, 15, and 20
Applicant argues that “the other cited references do not cure the deficiencies of Baeuml and Huang” with respect to dependent claims 6-7, 15, and 20. The examiner agrees that the prior combination does not teach the amended limitations incorporated into claims 1, 8, and 16. However, these deficiencies are addressed by the new rejection over Qiu and Huang and further in view of Cronin. Accordingly, Applicant’s arguments regarding the independent claims, and the dependent claims, are not persuasive against the new ground of rejection.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-4 are rejected under 35 U.S.C. 103 as being unpatentable over Qiu (US 12670899 B1) in view of Cronin et al. (US 20180276770 A1) and further in view of Kachuee et al. (US 12586575 B1).
Regarding claim 1, Qiu teaches a system for creating a storytelling experience (Col. 9, Ln. 4-6, “FIG. 2 illustrates a system 100 for using one or more language models to determine an action responsive to a user input.”), a microphone configured to collect data representative of speech (Col. 52, 56-58, “The user device 210 may include audio capture component(s), such as a microphone or array of microphones of a user device”) A display configured to display visualizations (Col. 68, Ln. 5-6, “The user device 210 may additionally include a display 1116 for displaying content.”) a speaker configured to play audio (Col. 67, Ln. 61-62, such as an audio output component such as a speaker”) a computing device comprising; processing circuitry (Col. 67, Ln. 20-22, “Each of these devices (110/120/125) may include one or more controllers/processors (1104/1204), which may each include a central processing unit (CPU)”); Memory, accessible by the processing circuitry and storing instructions that, when executed by the processing circuitry, cause the processing circuitry to perform operations (Col. 67, Ln. 28-30, “Each device (110/120/125) may also include a data storage component (1108/1208) for storing data and controller/processor-executable instructions.”) receiving, from the microphone, the data representative of the speech (Col. 62, Ln. 28-30, “the system component(s) may receive the audio data 811 from the user device 210, to recognize speech corresponding to a spoken input”) performing natural language understanding (NLU) on the received data to determine a semantic meaning of the speech (Col. 67, Ln. 10-15, “system components (120/125) may be included … such as one or more natural language processing system component(s) 120 for performing ASR processing, one or more natural language processing system component(s) 120 for performing NLU processing” where Col. 2, Ln. 11-14, “(NLU) is a field of computer science, artificial intelligence, and linguistics concerned with enabling computers to derive meaning from natural language inputs (such as spoken inputs)”); and receiving media wherein the media can comprise images. (Col. 10, Ln. 61-63, “the user input data 227 may correspond to various data types, such … image”) generating, via a large language model (LLM), based on the received data, the determined semantic meaning of the speech, and the media, a visualization and corresponding audio. (Col. 10, Ln. 12-14, “machine learning model(s) may receive text and/or other types of data as inputs (e.g., audio, image, video, etc.)”; Col. 14, Ln. 40-44, “The plan prompt generation component 310 processes the user input data 227 to generate prompt data 315 representing a prompt for input to the plan generation language model 320.”; Col. 22, Ln. 42-44, “user input data 227 (or a representation of a determined intent of the user input data 227”; Col. 48, Ln. 46-50, “generate output audio data including synthesized speech corresponding to the output data 262, which the system 100 (e.g., the response arbitration component 260.”. Thus, Qiu teaches that the user input data provided to the LLM can include speech, and input data, semantic meaning (intent), and generates a visualization with corresponding audio) may send to the user device 210 for output to the user 205.”) causing the visualization generated by the LLM to be displayed in real time via the display (Col. 48, Ln. 51-53, “generate visual output data (e.g., output image and/or video data) corresponding to the output data”. Under the broadest reasonable interpretation of ‘real time’ in light of the specification, a system receiving an input and rendering an output sequentially can be considered generating in real-time); and causing the audio generated by the LLM to be played in real-time via the speaker (Col. 48, Ln. 46-48, “generate output audio data including synthesized speech corresponding to the output data”. Under the broadest reasonable interpretation of ‘real time’ in light of the specification, a system receiving an input and rendering an output sequentially can be considered generating in real-time).
Qiu does not teach wherein the media comprises one or more images of a guest during a visit to an amusement park.
However, Cronin teaches wherein the media comprises one or more images of a guest during a visit to an amusement park. (Para 0007, “the stored images of a given patron's activity and the information stored in the theme park database associated with the given patron”).
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Qiu in order to incorporate the teachings of Cronin in order to generate more personalized content reflecting the guest’s visit (Para 0007).
Qiu modified Cronin does not teach where the output is rendered as the microphone collects the data representative of the speech.
However, Kachuee teaches where the output is rendered as the microphone collects the data representative of the speech. (Col. 14 Ln 59-62, “a user expressing a particular sentiment or emotion during output of the system response (e.g., sentiment detected from gestures and/or facial expressions, sentiment detected from voice)”).
It would have been obvious to one of ordinary skill in the art to modify Qiu before the effective filing date in order to incorporate the teachings of Kachuee in order to make the system more responsive by allowing the user to provide feedback or additional instructions without waiting for the output to finish. (Col. 14).
Regarding claim 2, Qiu teaches an imaging system comprising an imaging sensor configured to collect imaging data, wherein the operations comprise: receiving, from the imaging system, the imaging data (Col. 53, Ln. 6-8, “The user device 210 may also capture images using camera(s) of the user device 210”); and performing a gesture analysis on the received imaging data to identify one or more gestures captured in the imaging data (Col. 11, Ln. 24-27, “the user input may correspond … image data of a gesture”); wherein the visualization and the corresponding audio are generated by the LLM based on the imaging data, the one or more identified gestures, or any combination thereof. (Col. 14, Ln. 40-43, “The plan prompt generation component 310 processes the user input data 227 to generate prompt data 315 representing a prompt for input to the plan generation language model”).
Regarding claim 3, Qiu teaches generating, via the LLM, a design for a keepsake to be provided to the guest. (Col. 10, Ln. 24-25, “the output may be another type of data, such as audio, image, video, etc” where an image can comprise a design which can be applied to a keepsake).
Regarding claim 4, Qiu teaches wherein the keepsake comprises a video, an image, a sticker, a calendar, a book, a shirt, a hat, a pin, or any combination thereof (Col. 10, Ln. 24-25, “the output may be another type of data, such as audio, image, video, etc” where an image can comprise a design which can be applied to a keepsake).
Claims 5, 8-13, 16-19, and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Qiu (US 12670899 B1) in view of Huang et al. (US 20200126584), Cronin et al. (US 20180276770 A1) in view of Kachuee et al. (US 12586575 B1).
Regarding claim 5, Qiu as above in claim 1, does not teach the performing of a sentiment analysis on the received data to determining a sentiment of the speech.
However, Huang teaches the performing of a sentiment analysis on the received data to determining a sentiment of the speech (Para 0003: “…one or more sentiments…”).
It would be obvious to a person of ordinary skill in the art before the effective filing date to modify Qiu to incorporate the teachings of Huang before the effective filing date to determine sentiment data in order to achieve more accurate data for customer satisfaction and user feedback (Para 0050: “…sentiment information can describe…”).
Regarding claim 8 Qiu teaches receiving, via a microphone, data representative of speech (Col. 52, 56-58, “The user device 210 may include audio capture component(s), such as a microphone or array of microphones of a user device”); performing natural language understanding (NLU) on the received data to determine a semantic meaning of the speech (Col. 67, Ln. 10-15, “system components (120/125) may be included … such as one or more natural language processing system component(s) 120 for performing ASR processing, one or more natural language processing system component(s) 120 for performing NLU processing” where Col. 2, Ln. 11-14, “(NLU) is a field of computer science, artificial intelligence, and linguistics concerned with enabling computers to derive meaning from natural language inputs (such as spoken inputs)”); receiving media, wherein the media comprises one or more images (Col. 10, Ln. 12-14, “machine learning model(s) may receive text and/or other types of data as inputs (e.g., audio, image, video, etc.)”). generating, via a large language model (LLM), based on the received data, the determined semantic meaning of the speech, and the media, a visualization and corresponding audio (Col. 10, Ln. 12-14, “machine learning model(s) may receive text and/or other types of data as inputs (e.g., audio, image, video, etc.)”; Col. 14, Ln. 40-44, “The plan prompt generation component 310 processes the user input data 227 to generate prompt data 315 representing a prompt for input to the plan generation language model 320.”; Col. 22, Ln. 42-44, “user input data 227 (or a representation of a determined intent of the user input data 227”; Col. 48, Ln. 46-50, “generate output audio data including synthesized speech corresponding to the output data 262, which the system 100 (e.g., the response arbitration component 260.”. Thus, Qiu teaches that the user input data provided to the LLM can include speech, and input data, semantic meaning (intent), and generates a visualization with corresponding audio); causing the audio generated by the LLM to be played in real-time via a speaker (Col. 48, Ln. 46-48, “generate output audio data including synthesized speech corresponding to the output data”. Under the broadest reasonable interpretation of ‘real time’ in light of the specification, a system receiving an input and rendering an output sequentially can be considered generating in real-time); and causing the visualization generated by the LLM to be displayed in real-time (Col. 48, Ln. 51-53, “generate visual output data (e.g., output image and/or video data) corresponding to the output data”. Under the broadest reasonable interpretation of ‘real time’ in light of the specification, a system receiving an input and rendering an output sequentially can be considered generating in real-time)
Qiu does not teach performing a sentiment analysis on the received data to determine a sentiment of the speech; or wherein the input includes the determined sentiment of speech
However, Huang teaches performing a sentiment analysis on the received data to determine a sentiment of the speech (Abstract, “and sentiment information that conveys one or more sentiments associated with the audio content”); or wherein the input comprises the determined semantic meaning of the speech and the determined sentiment of the speech. (Abstract, “The technique generates the image(s) based on recognition of: semantic information that conveys one or more semantic topics associated with the audio content; and sentiment information that conveys one or more sentiments associated with the audio content.”)
It would be obvious to a person of ordinary skill in the art to modify Qiu to incorporate the teachings of Huang before the effective filing date to utilize the semantic meaning in the input to determine the sentiment of the text so that LLM can generate better results (Para 0054: “…extracts sentiment information…”); and to determine sentiment data in order to achieve more accurate data for customer satisfaction and user feedback (Para 0050: “…sentiment information can describe…”).
Qiu modified by Huang does not teach wherein the media comprises one or more images of a guest during a visit to an amusement park.
However, Cronin teaches wherein the media comprises one or more images of a guest during a visit to an amusement park. (Para 0007, “the stored images of a given patron's activity and the information stored in the theme park database associated with the given patron”).
It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Qiu in order to incorporate the teachings of Cronin in order to generate more personalized content reflecting the guest’s visit (Para 0007).
Qiu modified by Huang and Cronin does not teach where the output is rendered as the microphone collects the data representative of the speech.
However, Kachuee teaches where the output is rendered as the microphone collects the data representative of the speech. (Col. 14 Ln 59-62, “a user expressing a particular sentiment or emotion during output of the system response (e.g., sentiment detected from gestures and/or facial expressions, sentiment detected from voice)”).
It would have been obvious to one of ordinary skill in the art to modify Qiu before the effective filing date in order to incorporate the teachings of Kachuee in order to make the system more responsive by allowing the user to provide feedback or additional instructions without waiting for the output to finish. (Col. 14).
Regarding claims 9, Qiu as above in claim 8, teaches the medium wherein the NLU comprises applying one or more NLU algorithms or rule sets to the received data to determine the semantic meaning of the speech based on one or more words identified in the speech. (Col. 2, Ln. 11-14, “natural language understanding (NLU) is … concerned with enabling computers to derive meaning from natural language inputs *such as spoken inputs). ASR and NLU are often used together as part of a language processing component of a system.”; Col. 57, Ln. 21-27, “The ASR component 850 may transcribe the audio data 811 into text data” and “identify words that match the sequence of sounds of the speech represented in the audio data 811” thus, Qiu teaches applying ASR processing to identify words contained in received speech and applying NLU processing to derive the semantic meaning of the spoken input).
Regarding claim 10, Qiu does not teach the medium wherein the sentiment analysis comprises applying one or more sentiment-identifying algorithms or rule sets to the received data to determine the sentiment of the speech based on one or more words identified in the speech, tone of voice, intonation, or any combination thereof.
However, Huang teaches the medium wherein the sentiment analysis comprises applying one or more sentiment-identifying algorithms or rule sets to the received data to determine the sentiment of the speech based on one or more words identified in the speech, tone of voice, intonation, or any combination thereof. (Para 0004, “a sentiment classification engine for identifying sentiment information associated with the audio content”).
It would be obvious to a person of ordinary skill in the art to modify Qiu to incorporate the teachings of Huang before the effective filing date to utilize the semantic meaning in the input to determine the sentiment of the text so that LLM can generate better results (Para 0054: “…extracts sentiment information…”); and to determine sentiment data in order to achieve more accurate data for customer satisfaction and user feedback (Para 0050: “…sentiment information can describe…”).
Regarding claims 11, Qiu teaches the computer readable medium, wherein the received data representative of the speech comprises audio data (Col. 2, Ln. 42-43, “A system may receive a user input as speech. For example, a user may speak an input to a device. “).
Regarding claim 12, Qiu does not teach receiving data representative of speech in the form of a transcript.
However, Huang teaches receiving data representative of speech in the form of a transcript. (Huang [0049]: “A speech recognizer engine 204 converts the stream of audio features received from the preprocessing engine 202 to text information.”).
It would have been obvious to a person of ordinary skill in the art to modify Qiu before the effective filing date in such a way as to allow transcript inputs in order to gain the benefit generating images based on a text transcript as taught by Huang ([0052]: “…most closely match the input text…”).
Regarding claim 13, Qiu teaches receiving context data, wherein the context data is representative of one or more actions performed by a guest during a visit to an amusement park; wherein the visualization and the corresponding audio are generated by the LLM based on the context data. (Col. 14, Ln. 54-59, “contextual signals associated with a user that provided the user input, such as information associated with a user profile of the user (e.g., user ID, user behavioral information, user preferences, age, gender, historical user interaction data, devices associated with the user profile, etc.)” where behavioral information, historical user interactions can both represent one or more actions taken by a guest such as previously using the system).
Claim 16 recites substantially the same limitations as claim 8. Accordingly, it is rejected for the same reasons.
Claim 17 recites substantially the same limitations as claim 13. Accordingly, it is rejected for the same reasons.
Regarding claim 18, Qiu teaches receiving, from an imaging system, imaging data from the guest delivering the speech describing the experience (Col. 53, Ln. 6-8, “The user device 210 may also capture images using camera(s) of the user device 210); and performing a gesture analysis on the received imaging data to identify one or more gestures made by the guest (Col. 11, Ln. 24-27, “the user input may correspond … image data of a gesture”); wherein the visualization and the corresponding audio are generated by the LLM based on the imaging data, the one or more identified gestures, or any combination thereof. (Col. 14, Ln. 40-43, “The plan prompt generation component 310 processes the user input data 227 to generate prompt data 315 representing a prompt for input to the plan generation language model”).
Regarding claim 19, Qiu teaches generating, via the LLM, a design for a keepsake to be provided to the guest, wherein the keepsake comprises a video, an image, a sticker, a calendar, a book, a shirt, a hat, a pin, or any combination thereof (Col. 10, Ln. 24-25, “the output may be another type of data, such as audio, image, video, etc” where an image can comprise a design which can be applied to a keepsake).
Claim 21 recites substantially the same limitations as claim 19. Accordingly, it is rejected for the same reasons.
Claims 6, 7, 20 are rejected under 35 U.S.C. 103 as being unpatentable over Qiu (U.S. 20230343323) in view of Huang (U.S. 20200126584), Cronin (US 20180276770 A1) and Kachuee (US 12586575 B1) and further in view of Garvey et al. (U.S. 20240320591).
Regarding claim 6, Qiu does not teach generating guest satisfaction data based on the
received data, the determined semantic meaning of the speech, the determined sentiment of the
speech, the LLM, or any combination thereof.
However, Huang teaches the concept of determining semantic meaning of the speech, the
determined sentiment of the speech,
It would be obvious to a PHOSITA to modify Qiu to incorporate the teachings of Huang before the effective filing date to utilize the semantic meaning in the input determine the sentiment of the text so that LLM can generate better results ([0054]: “…extracts sentiment information…”).
Qiu modified by Huang does not teach generating guest satisfaction data based on the
received data.
However, Garvey teaches the system wherein guest satisfaction data can be generated
based on qualitative and quantitative data which under the broadest reasonable interpretation can
include speech, or the semantic meaning of speech. (Para 0004: “…qualitative and quantitative data...”). Garvey teaches that this can help to isolate or identify underperforming areas of a product and allow for design changes that are more accurate to improve the user experience.
It would be obvious to a person of ordinary skill int eh art before the effective filing date to modify Qiu in such a way as to incorporate these teachings to help to isolate or identify underperforming areas of a product and allow for design changes that are more accurate to improve the user experience. (Para 0004: “…qualitative and quantitative data...”).
Regarding claim 7, as in claim 6, Qiu does not teach a system wherein the guest satisfaction data identifies an aspect of a visit to an amusement park and an indication of a guest’s satisfaction with the aspect of the visit to amusement park indicated by the speech.
However, Garvey teaches the system wherein guest satisfaction data identifies an aspect of a visit as indicated by speech. (Para 0015: “In some embodiments, the systems...”). The satisfaction data can be useful for measuring and refining underperforming aspects of a product (Para 0004: “…product or service that are underperforming…”).
It would be obvious to a person of ordinary skill in the art before the effective filing date to modify Qiu before the effective filing data in such a way as to incorporate these teachings to help to isolate or identify underperforming areas of a product and allow for design changes that are more accurate to improve the user experience. (Para 0004: “…qualitative and quantitative data...”).
Claim 20 recites substantially the same limitations as claim 7. Accordingly, it is rejected for the same reasons.
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Qiu (U.S. 20230343323) in view of Huang (U.S. 20200126584), Cronin (US 20180276770 A1) and Kachuee (US 12586575 B1), as applied to claims 5, 8-13, 16-19, and 21 above, and further in view of Agarwal et al. (U.S. 20250310326).
Regarding claim 15, the combination of Qiu, Huang and Cronin teach images of a guest’s trip to an amusement park. (Cronin, para 0007, “the stored images of a given patron's activity and the information stored in the theme park database associated with the given patron”)
The combination of Qiu, Huang and Cronin do not teach wherein the one or more images were captured by a mobile device belonging to the guest.
However, Agarwal teaches wherein the one or more images were captured by a mobile device belonging to the guest. (Abstract: “…prompt the user to upload an image…”).
It would have been obvious to a person having ordinary skill in the art at the time of
the claimed invention to incorporate the teaching of Agarwal’s into the combination of Qiu and Huang because it would achieve the added benefit of being able to verify the user’s identity and ensure more stringent security (Abstract: “…the user's identity…”).
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL ALAN FOSTER JR. whose telephone number is (571)272-8874. The examiner can normally be reached M - F 8:00am - 6:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at (571) 272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL A FOSTER JR/Examiner, Art Unit 2654
/Richa Sonifrank/Primary Examiner, Art Unit 2654