DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Examiner’s Note
This Office Action is in response to application filed on 7/6/2023, where claims 1-20 are currently pending.
Claim Objections
Claim 19 is objected to because of the following informalities: claim 19 recites “identifying the name of the object from speech that is received from speech that is received from a meeting device”. The phrase “speech that is received from” is duplicated in the above limitation. It is suggested to remove the second iteration as it is unnecessary. Appropriate correction is required.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-7, 9, 11-15, 17, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Aggarwal et al., (US 2025/0259003 A1) (hereinafter Aggarwal) in view of Robert Jose et al., (US 2023/0124889 A1) (hereinafter Robert).
Referring to claim 1, Aggarwal an apparatus comprising:
a processor configured to
receive a natural language input submitted from a user device (¶ [0059], fig. 3, “user 305, while interacting with device 310, may utter speech 340, such as, ‘find me items related to nurnia’”),
identify a name of an object based on the natural language input (¶ [0059], fig. 3, “context aware correction model 325 may identify ‘nurnia’ as a candidate term for substitution.”),
map the name of the object to a name of an alternative object (¶ [0060], fig. 3, “context aware correction model 325 can access on-device repository 360 of pairs of mistranscribed terms and non-common terms. As an illustrative example, the pairs of mistranscribed terms and non-common terms can be ‘(haarnia, hernia),’ ‘(narnia, hernia),’ and ‘(hair near, hernia),’ where mistranscribed terms, ‘haarnia,’ ‘narnia,’ and ‘hair near’ are paired with a non-common term ‘hernia.’ As illustrated in block 355, context aware correction model 325 can compare the transcribed term ‘nurnia’ 350 to the mistranscribed terms, and determine that ‘nurnia’ 350 matches the mistranscribed term ‘narnia.’”) and identify a web page corresponding to the alternative object, via a rule set (¶ [0030], “the input recognition can be performed by a remote server. In such embodiments, the on-device system can receive the input and send the input to the remote server. The remote server can apply the first trained machine learning model to generate the transcription of the input. Subsequently, the on-device system can receive the transcription from the remote server, and perform the auto-correction of the transcription.” ¶ [0115], fig. 7, “Server devices 708, 710 can be configured to perform one or more services, as requested by programmable devices 704a-704e. For example, server device 708 and/or 710 can provide content to programmable devices 704a-704e. The content can include, but is not limited to, web pages”.)
Aggarwal teaches the limitations above. However, Aggarwal does not explicitly teach display an option to navigate to the identified web page that corresponds to the…object via a user interface displayed on the user device.
Robert teaches display an option to navigate to the identified web page that corresponds to the…object via a user interface displayed on the user device (¶ [0004], “a method for providing contextual based actions based on a natural language input. The method further comprises: receiving, on a media device, a natural language input; determining, based on the natural language input, a first context of the natural language input; determining, based on the first context, a first action…generating for display on the media device a dynamic action button configured to be selected by a user to carry out an action”. ¶ [0070], “the system may not order through applications automatically and require user input. In such examples, the dynamic action button 310 can be configured to be ‘deep-linked,’ i.e., clicking the dynamic action button 310 automatically launches the user's favourite food delivery app to the pizza food page.”)
Aggarwal and Robert are analogous art to the claimed invention because they are concerning with interface for receiving natural language input (i.e., same field of endeavor).
It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention having Aggarwal and Robert before them to modify the machine learning based context aware correction apparatus of Aggarwal to incorporate the function of generating a dynamic action button based on first context as taught by Robert. One of ordinary skill in the art would have combined the elements as claimed by known methods as disclosed by Robert (¶ [0064]-[0072]), because the function of generating a dynamic action button based on first context does not depend on the machine learning based context aware correction apparatus. That is the function of generating a dynamic action button based on first context performs the same function independent on which interface it is incorporated onto, and therefore, the result of the combination would have been predictable to one of ordinary skill in the art. The motivation to combine would have been to allow user to take appropriate action based on context of an input with ease as suggested by Robert (¶ [0018]).
Referring to claim 3, Aggarwal further teaches the apparatus of claim 1, wherein the processor is configured to
identify the name of the object from speech that is received from a meeting device of a teleconference (¶ [0156], “the interaction with the computing device may be one of an interaction with…a video communications application”).
Aggarwal teaches the limitations above. However, Aggarwal does not explicitly teach display a prompt via the meeting device which includes the option to navigate to the identified web page.
Robert further teaches display a prompt…which includes the option to navigate to the identified web page (¶ [0013], “generating for display on the media device a dynamic action button, configured to be selected by the user to carry out an action”).
Referring to claim 4, Aggarwal further teaches the apparatus of claim 1, wherein the processor configured is configured to
generate a navigation path based on web text on web pages included in the natural language input submitted from the user device, and
map the name of the object to the name of the alternative object based on the navigation path (¶ [0079], “The one or more past interactions of the user with the computing device can involve textual content provided by the viewer interface. For example, the user may be browsing documents on the web, browsing a library associated with media content, and so forth. The non-common term in the plurality of pairs appears can appear the textual content, such as the documents that are being browsed”).
Referring to claim 5, Aggarwal further teaches the apparatus of claim 1, wherein the processor is further configured to
receive browsing history from a browser on the user device, and
further identify…that corresponds to the alternative object based on the browsing history (¶ [0146], “the non-common terms in the plurality of pairs were observed by the on-device system in one or more past interactions of the user with the computing device.” ¶ [0079], “The one or more past interactions of the user with the computing device can involve textual content provided by the viewer interface. For example, the user may be browsing documents on the web”. ¶ [0151], “the one or more past interactions of the user with the computing device may be one of…a web browsing application”.)
Aggarwal teaches the limitations above. However, Aggarwal does not explicitly teach identify the web page.
Robert further teaches identify the web page (¶ [0070], “the system may not order through applications automatically and require user input. In such examples, the dynamic action button 310 can be configured to be ‘deep-linked,’ i.e., clicking the dynamic action button 310 automatically launches the user's favourite food delivery app to the pizza food page.”)
Referring to claim 6, Aggarwal further teaches the apparatus of claim 1, wherein the processor is configured to map the name of the object to the name of the alternative object via execution of a machine learning model on the natural language input (¶ [0021], “machine learning based context aware correction for input recognition (e.g., automatic speech recognition (ASR)). In particular, a trained machine learning model (e.g., a convolutional neural network) can recognize a potentially mistranscribed term in speech as transcribed by the ASR system, and substitute the potentially mistranscribed term with another term in the speech as transcribed.”)
Referring to claim 7, Aggarwal further teaches the apparatus of claim 6, wherein the processor is further configured to
detect whether the user selects…that corresponds to the alternative object (¶ [0076], “one or more of a determination of the mistranscribed terms, the pairs comprising a mistranscribed term, a confidence level associated with each pair, a threshold for acceptance of a pair for substitution, can vary from one user to another”. ¶ [0077], “user feedback may be incorporated in such determinations.”), and
retrain the machine learning model based on the detected user selection and the alternative object (¶ [0111], “Input data 630 can include a non-common term. Other types of input data are possible as well. Inference(s) and/or prediction(s) 650 can include one or more mistranscribed terms that represent phonetically similar alternatives of pronouncing the non-common term. In some embodiments, inference(s) and/or prediction(s) 650 can include pairs of non-common terms and mistranscribed terms, with an associated confidence level indicative of a phonetic similarity of the mistranscribed term and the non-common term. Inference(s) and/or prediction(s) 650 can include other output data produced by trained machine learning model(s) 632 operating on input data 630 (and training data 610). In some examples, trained machine learning model(s) 632 can use output inference(s) and/or prediction(s) 650 as input feedback 660.”)
Aggarwal teaches the limitations above. However, Aggarwal does not explicitly teach the option to navigate to the identified web page.
Robert further teaches the option to navigate to the identified web page (¶ [0070], “the system may not order through applications automatically and require user input. In such examples, the dynamic action button 310 can be configured to be ‘deep-linked,’ i.e., clicking the dynamic action button 310 automatically launches the user's favourite food delivery app to the pizza food page.”)
Regarding claims 9 and 11-15, these claims recite the method performed by the apparatus of claims 1 and 3-7 respectively; therefore, the same rationale of rejection is applicable.
Regarding claims 17, 19, and 20, these claims recite the computer-program product comprising a computer readable storage medium having stored thereon instructions when executed by a processor to perform the same method steps performed by the apparatus of claims 1, 3, and 4 respectively; therefore, the same rationale of rejection is applicable.
Claims 2, 10, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Aggarwal in view of Robert as applied to claims 1, 9, and 17 above, and further in view of Wang, (US 10,379,696 B2) (hereinafter Wang).
Referring to claim 2, Aggarwal further teaches the apparatus of claim 1, wherein the processor is configured to
identify the name of the object from a web page where the object is visible within a browser on the user device (¶ [0038], “user interaction may be an interaction with a search assistant. For example, the user may input text into a search field of a web browser.”).
Aggarwal teaches the limitations above. However, Aggarwal does not explicitly teach overlay a pop-up window within the browser where the object is visible which includes the option to navigate to the identified web page.
Robert further teaches within the browser where the object is visible which includes the option to navigate to the identified web page (¶ [0070], “the system may not order through applications automatically and require user input. In such examples, the dynamic action button 310 can be configured to be ‘deep-linked,’ i.e., clicking the dynamic action button 310 automatically launches the user's favourite food delivery app to the pizza food page.”)
Aggarwal in view of Robert teach the limitations above. However, Aggarwal in view of Robert do not explicitly teach a pop-up window.
Wang teaches a pop-up window (Figure 2-2 shows a popup window with buttons).
Aggarwal, Robert, and Wang are analogous art to the claimed invention because they are concerning with interface for allowing user interaction (i.e., same field of endeavor).
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention having Aggarwal in view of Robert and Wang before them to substitute the pop-up window as taught by Wang for the generic interface of Aggarwal in view of Robert. Because both Aggarwal in view of Robert and Wang teach methods of presenting user interface, it would have been obvious to one skilled in the art to substitute one known method for the other to achieve the predictable result of graphical user interface technology. The motivation would have been to provide additional function without rearranging the original content.
Regarding claim 10, the instant claim recites the method performed by the apparatus of claim 2; therefore, the same rationale of rejection is applicable.
Regarding claim 18, the instant claim recites the computer-program product comprising a computer readable storage medium having stored thereon instructions when executed by a processor to perform the same method steps performed by the apparatus of claim 2; therefore, the same rationale of rejection is applicable.
Claims 8 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Aggarwal in view of Robert as applied to claims 1 and 9 above, and further in view of Mariani et al., (US 2011/0298799 A1) (hereinafter Mariani).
Referring to claim 8, Aggarwal teaches the apparatus of claim 1. However, Aggarwal does not explicitly teach wherein the processor is configured to identify an image included in the natural language input and modify the image to generate a new image based on the alternative object.
Robert further teaches identify an image included in the natural language input (¶ [0112], “the contextual factors that accompany the user input includes sensor information, e.g., lighting, ambient noise, ambient temperature, images”).
Aggarwal in view of Robert teach the limitations above. However, Aggarwal in view of Robert do not explicitly teach modify the image to generate a new image based on the alternative object.
Mariani teaches modify the image to generate a new image based on the alternative object (¶ [0008], “a method and a system for replacing a first object in a 2D image with a second object based on a synthesized three-dimensional (3D) model of the second object.”)
Aggarwal, Robert, and Mariani are analogous art to the claimed invention because they are concerning with interface for manipulating input data (i.e., same field of endeavor).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention having Aggarwal in view of Robert and Mariani before them to modify the machine learning based context aware correction apparatus of Aggarwal in view of Robert to incorporate the function of generating a new image based on an alternative object by Mariani. One of ordinary skill in the art would have combined the elements as claimed by known methods as disclosed by Mariani (¶ [0008]-[0010]), because the function of generating a new image based on an alternative object does not depend on the machine learning based context aware correction apparatus. That is the function of generating a new image based on an alternative object performs the same function independent on which interface it is incorporated onto, and therefore, the result of the combination would have been predictable to one of ordinary skill in the art. The motivation to combine would have been to provide more desirable performance in detecting an object in an image as suggested by Mariani (¶ [0007]).
Regarding claim 16, the instant claim recites the method performed by the apparatus of claim 8; therefore, the same rationale of rejection is applicable.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
US 2023/0289530 (Aung) – discloses menu ordering system that interprets natural language input.
US 2017/0160813 (Divakaran) – discloses virtual personal assistant that accepts natural language input.
US 2015/0371632 (Skobeltsyn) – discloses methods, systems, and apparatus for recognizing names of entities in speech.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MONG-SHUNE CHUNG whose telephone number is (571) 270-5817. The examiner can normally be reached on M-F (9-5) EST.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Scott Baderman, can be reached at telephone number 571-272-3644. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from Patent Center and the Private Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from Patent Center or Private PAIR. Status information for unpublished applications is available through Patent Center and Private PAIR for authorized users only. Should you have questions about access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) Form at https://www.uspto.gov/patents/uspto-automated- interview-request-air-form.
/MONG-SHUNE CHUNG/
Primary Examiner, Art Unit 2118