DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
This is in response to applicant’s amendment/response filed on 08/05/2026, which has been entered and made of record. Claims 1, 8, and 15 have been amended. No Claim has been cancelled. No Claim has been added. Claims 1-20 are pending in the application.
Response to Arguments
Applicant's arguments filed 08/05/2026 have been fully considered but they are not persuasive.
Applicant submits that “Moreau does not disclose or suggest "determining second attributes of the action, the second attributes comprising a type of the action," much less "ranking plural augmented reality content items based on the assigned weights and on the second attributes of the action, the ranking comprising comparing the first attributes of the object and the second attributes of the action with predefined attributes associated with each of the plural augmented reality content items." (Remarks, p. 9).
The examiner disagrees with Applicant’s premises and conclusion.
Moreau recites “those identified product features are ranked based on the type of user interaction with each corresponding portion. For instance, product features corresponding with portions of a video or audio the recipient viewed, replayed, paused on, zoomed, or otherwise positively engaged with are given a higher ranking Conversely, product features corresponding with portions of a video or audio the user skipped or other otherwise negatively engaged with are given a lower ranking.” (¶35). In other words, Moreau indeed teaches determining a type of the action and ranking the product features.
The arguments regarding independent claims 8 and 15 are moot for at least the reasons discussed above.
The arguments regarding dependent claims for the virtue of their dependency are moot because the independent claims are not allowable.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 8-9, and 15-16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Cowburn et al. (US 10074381 B1), in view of Manjunath et al. (US 20210089124 A1), and further in view of Moreau et al. (US 20160330155 A1).
Regarding Claim 8, Cowburn discloses A device (ABST reciting “an augmented reality system ”. Fig. 15 showing a machine.), comprising:
at least one processor; (Fig. 15 showing processors 1504) and
a memory (Fig. 15 showing memory/storage 1506) storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations (col. 21, ln. 20-35 reciting “The machine 1500 may comprise, . . ., or any machine capable of executing the instructions 1510, sequentially or otherwise, that specify actions to be taken by machine 1500. Further, while only a single machine 1500 is illustrated, the term “machine” shall also be taken to include a collection of machines that individually or jointly execute the instructions 1510 to perform any one or more of the methodologies discussed herein.”) comprising:
causing, by a messaging application running on the device, a camera of the device to capture an image; (col. 9, ln. 47-50 reciting “A message image payload 406: image data, captured by a camera component of a client device 102 . . ., and that is included in the message 400.”)
receiving, by the messaging application, speech input to select augmented reality content for display with the image; (ABST reciting “Various embodiments may detect speech, identify a source of the speech, transcribe the speech to a text string, generate a speech bubble based on properties of the speech and that includes a presentation of the text string, and cause display of the speech bubble at a location in the augmented reality interface based on the source of the speech.” Further, col. 19, ln. 15-18 reciting “As discussed in operation 804 of FIG. 8, the augmented reality system 150 may select a speech bubble from a speech bubble library based on an emotional effect of the detected speech and the speech properties.”)
determining at least one keyword included in the speech input; (col. 3, ln. 43-45 reciting “the augmented reality system may parse the text string of the transcribed speech into individual words”)
determining that the at least one keyword indicates an action to perform with respect to the image; (col. 3, ln. 43-49 reciting “the augmented reality system may parse the text string of the transcribed speech into individual words, determine a definition of the set of words, and compare the definitions of the words to an emotional effect library. Based on the comparison, the augmented reality system may determine an intended emotional effect of the speech.” Further, col. 16, ln. 27-37 reciting “The speech bubble themes indicate a design and form to be applied to the speech bubble based on the emotional effect. For example, an emotional effect of “angry” may have a corresponding speech bubble theme that causes the speech bubble to display as a red jagged bubble, with red text and animated fire, while an emotional effect of “sad” may have a corresponding speech bubble theme that causes the speech bubble to display as a drooping blue bubble with frowny faces and black text”)
determining first attributes of the object depicted in the image; (col. 2, ln. 65-67 reciting “the presentation of the space may include depictions of multiple people, each with corresponding user profiles.”, where the user profiles correspond to first attributes of the depicted people. In addition, col. 3, ln. 63-67 reciting “having identifies a user as the source of the speech, the augmented reality system may capture facial landmarks of the user and apply facial landmark recognition techniques to determine an emotional state of the user (e.g., based on a smile, a frown, a furrowed brow, etc.).”, where the user corresponds to an object depicted in the image, and the facial landmark corresponds to the first attributes. )
However, Cowburn does not explicitly disclose the at least one keyword indicates an object depicted in the image.
Manjunath teaches “electronic systems and techniques for using such systems in relation to various simulated reality technologies” (¶10). More specifically, ¶40 recites “For physical objects in SR setting 300, image data of physical setting 306 is obtained from one or more second image sensors 214b to identify the physical objects and determine corresponding attribute tags for the physical objects. For example, computer vision module 220 obtains image data of physical setting 306 from one or more second image sensors 214b (via connection 208) and performs pattern recognition to identify physical objects 308-318. As discussed above, the corresponding attribute tags are stored in association with unique physical object identifiers that are assigned by reality engine 218 to each of physical objects 308-318.”; and further ¶55 recites “In another illustrative example where the speech input is “What model of laptop is that,” natural language understanding module 236 determines that the domain corresponding to the text representation of this speech input is the “search” domain. . . In particular, the parameter resolution module 238 determines that the word “that” in the speech input is referring to physical object 308 and thus the “search object(s)” parameter is resolved as including an image of physical object 308 and the text search strings “model” and “laptop.”” Thus Manjunath teaches the at least one keyword indicates an object depicted in the image.
It would have been obvious to one with ordinary skill, before the effective filing date of the claimed invention, to modify the device (taught by Cowburn) to determine at least one keyword that indicates an object depicted in the image (taught by Manjunath). The suggestions/motivations would have been that “As a result, user experience is enhanced, which corresponds to improved operability of the voice assistant operating on the electronic system.” (¶4), and to apply a known technique to a known device (method, or product) ready for improvement to yield predictable results.
Cowburn discloses activating an augmented reality content item (i.e. choosing a bubble from the bubble library) based on the attribute of the object (i.e. the talking person’s facial landmarks). However, Cowburn in view of Manjunath does not explicitly disclose
assigning weights to each of the first attributes of the object;
determining second attributes of the action, the second attributes comprising a type of the action;
ranking plural augmented reality content items based on the assigned weights and on the second attributes of the action, the ranking comprising comparing the first attributes of the object and the second attributes of the action with predefined attributes associated with each of the plural augmented reality content items;
selecting, based on the ranking, a highest-ranked augmented reality content item from among the plural augmented reality content items; and
activating the highest-ranked augmented reality content item with respect to the image.
Moreau teaches content ranking method based on weight assign to features, and “product feature rankings are tracked by assigning a weight to each product feature” (¶39) . More specifically, Moreau teaches a highest-ranked item is chosen based on the weight of each product feature, and recites “based on the recipient's interaction with another rich media component, the product feature rankings indicate a highest ranking for the Palace of Versailles, . . . Based on the product feature rankings, when the recipient accesses the link to the video, the video is streamed such that the portion from 12-26 seconds (corresponding to the Palace of Versailles) is streamed first” (¶41), “a highest ranking for the Palace of Versailles, . . ., an image collection may be generated by providing images of the Palace of Versailles first,” (¶42). In addition, ¶35 teaches ranking product features based on a type of the user action, and recites “In the case of user interactions with a streaming video or audio, product features associated with each portion with which the recipient interacted are identified based on timeline tagging or sub-video/sub-audio tagging, and those identified product features are ranked based on the type of user interaction with each corresponding portion. For instance, product features corresponding with portions of a video or audio the recipient viewed, replayed, paused on, zoomed, or otherwise positively engaged with are given a higher ranking Conversely, product features corresponding with portions of a video or audio the user skipped or other otherwise negatively engaged with are given a lower ranking.”
It would have been obvious to one with ordinary skill, before the effective filing date of the claimed invention, to modify the method of the augmented item ranking based on object attributes (taught by Cowburn in view of Manjunath) to weight the attributes of the object and to rank based on the weight and a type of the user action (taught by Moreau). The suggestions/motivations would have been that “This provides a highly relevant and personalized experience to the recipient that may be accomplished irrespective of the order in which the recipient interacts with the various rich media components.” (¶4), and to apply a known technique to a known device (method, or product) ready for improvement to yield predictable results.
Regarding Claim 9, Cowburn in view of Manjunath and Moreau discloses The device of claim 8, the operations further comprising:
performing a scan of the image to identify multiple objects in the image; and
detecting, based on performing the scan, the object from among the multiple objects.
(Cowburn, col. 2, ln. 65 – col. 3, ln. 7 reciting “For example, the presentation of the space may include depictions of multiple people, each with corresponding user profiles. Having received the speech data (e.g., the speech recorded through the microphone), the augmented reality system identifies a user profile based on the speech data through speech recognition techniques. Upon determining the user profile based on the speech data, the augmented reality system determines which individual depicted in the presentation is the source of the speech based on facial landmark recognition data.”)
Claim 1, has similar limitations as of Claim(s) 8, therefore it is rejected under the same rationale as Claim(s) 8.
Claim 2, has similar limitations as of Claim(s) 9, therefore it is rejected under the same rationale as Claim(s) 9.
Claim 15, has similar limitations as of Claim(s) 8, therefore it is rejected under the same rationale as Claim(s) 8.
Claim 16, has similar limitations as of Claim(s) 9 and 15, therefore it is rejected under the same rationale as Claim(s) 9 and 15.
Claim(s) 3, 5, 10, 12, 17 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Cowburn et al. (US 10074381 B1) in view of Manjunath and Moreau et al. (US 20160330155 A1), and further in view of Xie et al. (US 20190235833 A1).
Regarding Claim 10, Cowburn in view of Manjunath and Moreau discloses The device of claim 8.
Cowburn discloses performing speech recognition based on the speech input and receiving the at least one keyword, and recites “the augmented reality system may parse the text string of the transcribed speech into individual words” (col. 3, ln. 43-45). However, Cowburn in view of Manjunath and Moreau does not explicitly disclose wherein determining the at least one keyword comprises:
sending, to a speech recognition service, a request to perform speech recognition based on the speech input; and
receiving, from the speech recognition service and based on sending the request, the at least one keyword.
Xie teaches “a method and system based on speech and augmented reality environment interaction.” (ABST). More specifically, Xie teaches “to invoke Automatic Speech Recognition (ASR) service to parse the user's speech data to obtain a speech recognition result corresponding to the speech. The speech recognition result is a recognized text corresponding to the speech.” (¶53), In addition, ¶85 recites “to invoke Automatic Speech Recognition (ASR) service to parse the user's speech data to obtain a speech recognition result corresponding to the speech. The speech recognition result is a recognized text corresponding to the speech.”
It would have been obvious to one with ordinary skill, before the effective filing date of the claimed invention, to modify the device (taught by Cowburn in view of Manjunath and Moreau) to invoke a speech recognition service to receive at least one keyword (taught by Xie). The suggestions/motivations would have been to apply a known technique to a known device (method, or product) ready for improvement to yield predictable results.
Claim 3, has similar limitations as of Claim(s) 10, therefore it is rejected under the same rationale as Claim(s) 10.
Claim 17, has similar limitations as of Claim(s) 10, therefore it is rejected under the same rationale as Claim(s) 10.
Regarding Claim 12, Cowburn in view of Manjunath, Moreau and Xie discloses The device of claim 8, where the at least one keyword comprises a first keyword indicating the object and a second keyword indicating the action to perform with respect to the object. (Xie, ¶51 disclosing an example voice command having two keywords, and reciting “prompts such as “rotate the model”, “enlarge the model” and “reduce the model” are displayed in the scenario, the user may input a formatted fixed speech according to the above prompts”. The suggestions/motivations would have been the same as that of Claim 10 rejections.)
Claim 5, has similar limitations as of Claim(s) 12, therefore it is rejected under the same rationale as Claim(s) 12.
Claim 19, has similar limitations as of Claim(s) 12 and 15, therefore it is rejected under the same rationale as Claim(s) 12 and 15.
Claim(s) 4, 11, and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Cowburn et al. (US 10074381 B1), in view of Manjunath and Moreau et al. (US 20160330155 A1), and further in view of Xie et al. (US 20190235833 A1), and further in view of Rathod (WO 2018104834 A1).
Regarding Claim 11, Cowburn in view of Manjunath, Moreau and Xie discloses The device of claim 10.
However, Cowburn in view of Manjunath, Moreau and Xie does not explicitly disclose wherein a first part of the speech input includes a trigger word, and
wherein the at least one keyword is based on a second part of the speech input that does not include the trigger word.
Rathod teaches “present invention enables user to provide voice command to start video talk with voice command related contact” (p. 20, ln. 11-30). In other words, Rathod teaches a voice command including a trigger word to start an application.
Xie discloses “it is possible to monitor the user's speech data after activating the augmented reality function” (¶50). Further, ¶51 recites “if a current scenario is a preset augmented reality sub-environment scenario, it is possible to guide the user to input a preset speech operation instruction. . . , the user enters the preset augmented reality sub-environment scenario by clicking a specific entrance, and the vehicle 3D model is displayed in the preset augmented reality sub-environment scenario.”
It would have been obvious to one with ordinary skill, before the effective filing date of the claimed invention, to modify the device (taught by Cowburn in view of Manjunath, Moreau and Xie) to start an application, such activating an augmented reality function, with a voice command including a trigger word (taught by Rathod). The suggestions/motivations would have been “to identify user intention to take photo or video and automatically invokes, open and show camera display screen so user is enable to capture photo or video without each time manually open camera application.” (p. 21, ln. 3-5), and to apply a known technique to a known device (method, or product) ready for improvement to yield predictable results.
Claim 4, has similar limitations as of Claim(s) 11, therefore it is rejected under the same rationale as Claim(s) 11.
Claim 18, has similar limitations as of Claim(s) 11, therefore it is rejected under the same rationale as Claim(s) 11.
Claim(s) 6-7, 13-14, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Cowburn et al. (US 10074381 B1), in view of Manjunath and Moreau et al. (US 20160330155 A1), and further in view of Sherr et al. (US 20020154157 A1).
Regarding Claim 13, Cowburn in view of Manjunath and Moreau discloses The device of claim 8.
Cowburn in view of Manjunath and Moreau does not explicitly disclose the operations further comprising:
displaying, in ranked order based on the ranking, an interface with user-selectable elements for activating remaining ones of the plural augmented reality content items.
Sherr teaches selecting movies by title for display and displaying a carousel interface with user-selectable elements for activating a content, and recites “if the method of organizing the browse feature is by title and the display method is by virtual carousel, the content items (for example, movies) may then be displayed by title in alphabetical order in a virtual carousel interface. In FIG. 5, the shaded icons or areas in the menus 502 and 504 represent a state in which the user has selected to browse content item information (such as movie information) 508 by title and to display representations of the content items (for example, box art from movie video boxes) in a virtual carousel display, i.e., in horizontal rows 516, which may be caused to appear to move or spin to the left or right by selecting icon or area 510 or icon or area 512, respectively, simulating the motion of a carousel. A user may interact with the display representations of content items presented on the Browse page in the same way as content item representations on the home page, for example, by selecting a content item (for example, movie) with a left or right click to view more information or purchase or obtain a license to access the selected content item, respectively.” (¶85).
It would have been obvious to one with ordinary skill, before the effective filing date of the claimed invention, to apply the device (taught by Cowburn in view of Manjunath and Moreau) to select contents displayed on an interface based on an order of the object attributes (movie titles) and action (playing a movie) (taught by Sherr). The suggestions/motivations would have been that “there is a demand for user interfaces, . . ., which are not only easy to operate, but also provide a distinguishable format, an opportunity to obtain various types of information about content pieces, and an inducement to select content pieces.” (¶8), and to apply a known technique to a known device (method, or product) ready for improvement to yield predictable results.
Regarding Claim 14, Cowburn in view of Manjunath, Moreau and Sherr discloses The device of claim 13,
wherein the interface is a carousel interface with a respective user-selectable icon for each of the plural augmented reality content items, and
wherein the carousel interface differentiates display of the icon for the augmented reality content item, relative to remaining icons, within the carousel interface.
(Sherr, Fig. 5. ¶85 reciting “In FIG. 5, the shaded icons or areas in the menus 502 and 504 represent a state in which the user has selected to browse content item information (such as movie information) 508 by title and to display representations of the content items (for example, box art from movie video boxes) in a virtual carousel display, i.e., in horizontal rows 516, which may be caused to appear to move or spin to the left or right by selecting icon or area 510 or icon or area 512, respectively, simulating the motion of a carousel. A user may interact with the display representations of content items”. The suggestions/motivations would have been the same as that of Claim 13 rejections.)
Claim 6, has similar limitations as of Claim(s) 13, therefore it is rejected under the same rationale as Claim(s) 13.
Claim 7, has similar limitations as of Claim(s) 14, therefore it is rejected under the same rationale as Claim(s) 14.
Claim 20, has similar limitations as of Claim(s) 13 and 15, therefore it is rejected under the same rationale as Claim(s) 13 and 15.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YI WANG whose telephone number is (571)272-6022. The examiner can normally be reached 9am - 5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jason Chan can be reached at (571)272-3022. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YI WANG/Primary Examiner, Art Unit 2619