DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 02/13/2025, 05/15/2025, 12/29/2025 and 05/01/2026 were filed. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-10, 12 and 15-23 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The claims are generally narrative and indefinite, failing to conform with current U.S. practice. They appear to be a literal translation into English from a foreign document and are replete with grammatical and idiomatic errors.
Regarding claims 1, 12 and 23, Independent claims 1, 12 and 23 respectively recite “obtaining, in response to determining that a first image of a physical space in which the robotic device is located indicates that the physical space comprises the first object and the second object, the first object and the second object by the robotic device”. Examiner submits that in the current form, the claim language recites that the obtaining step occurs “in response to determining that a first image of a physical space in which the robotic device is located indicates that the physical space comprises the first object and the second object”. Examiner submits that it is unclear if the obtaining step is executed only in response to the determination of the first image. Therefore, the scope of the claims as currently presented cannot be properly ascertained by the examiner, thus the claim is unclear and indefinite. Appropriate correction and/or clarification is earnestly solicited.
Examiner further notes for examination purposes, the above mentioned limitation will be interpreted as obtaining a first and second object that are found within a robotic devices physical space based on imaging data.
Examiner notes a possible example to correct the 112(b) issue: “determine that a first image of a robotic device's physical space indicates the presence of a first object and a second object; in response to the determination of the first image, the robotic device obtains the first and second objects.”.
Regarding claims 2-10 and 15-22, Claims are rejected based on their dependency to a rejected claim.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-6, 12, 15-19 and 23 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Nakamura (US 2024/0278424 A1).
Regarding claim 1, Nakamura teaches a method for performing a user task, comprising: receiving, by a robotic device, a user task from a user, the user task instructing the robotic device to obtain a first object; determining a second object associated with the first object [(see at least paragraphs 37-39, 47-48) As in 37 “The robot control system 100 includes a robot 2 and a robot control device 110. The robot 2 moves a work target object 8 from a work start point 6 to a work target point 7 by means of an end effector 2B (grasping hand) at the end of an arm 2A. That is, the robot control device 110 controls the robot 2 so that the work target object 8 moves from the work start point 6 to the work target point 7. The robot control system 100 includes a camera 4. The camera 4 captures images of items and the like in an influence zone 5 that may affect operation of the robot 2.” As in 39 “the library group preferably includes many libraries so that many types of the work target object 8 may be recognized in the robot control system 100. However, when many libraries are present, selecting an appropriate library for the type of the work target object 8 may be difficult for a user. Further, before incorporating the robot control system 100, a user may want to know whether a target object can be properly recognized using a library in the library group. Further, even after the robot control system 100 has been incorporated, a user may want to know whether a new work target object 8 can be recognized.”]; and obtaining, in response to determining that a first image of a physical space in which the robotic device is located indicates that the physical space comprises the first object and the second object, the first object and the second object by the robotic device. [(see at least paragraph 37-40,48) As in 38 “The robot control device 110 recognizes the work target object 8 present in a space where the robot 2 works based on images captured by the camera 4. Recognition of the work target object 8 is executed using a library (trained model generated to recognize the work target object 8) included in the library group. The robot control device 110 downloads and deploys a library via the network 40 before executing recognition of the work target object 8.” As in 37 “FIG. 2 is a schematic diagram illustrating an example configuration of the robot control system 100. The robot control system 100 includes a robot 2 and a robot control device 110. The robot 2 moves a work target object 8 from a work start point 6 to a work target point 7 by means of an end effector 2B (grasping hand) at the end of an arm 2A. That is, the robot control device 110 controls the robot 2 so that the work target object 8 moves from the work start point 6 to the work target point 7. The robot control system 100 includes a camera 4. The camera 4 captures images of items and the like in an influence zone 5 that may affect operation of the robot 2.” As in 48 “According to the present embodiment, the defined task is recognition of the target object and display of a recognition result. For example, when the defined task is executed with respect to an image including multiple target objects (such as industrial components), recognition (identification of bolts, nuts, springs, and the like) is executed with respect to each of the multiple target objects, and the recognition result (names of bolts, nuts, springs, and the like) is overlaid on the image and output.”]
Regarding claim 2, Nakamura teaches wherein determining the second object comprises: obtaining a first prompt based on the user task and user information of the user, the first prompt configured to determine the second object; and receiving a first response of a machine learning model for the first prompt to determine the second object. [(see at least paragraphs 31-34, 48, 61-62) As in 31 “A user may download and use a library via a network 40, for example. According to the present embodiment, a library includes, but is not limited to, at least one machine learning model (trained model) generated by machine learning to enable a task such as described above to be executed.” As in 48 “For example, when the defined task is executed with respect to an image including multiple target objects (such as industrial components), recognition (identification of bolts, nuts, springs, and the like) is executed with respect to each of the multiple target objects, and the recognition result (names of bolts, nuts, springs, and the like) is overlaid on the image and output.”]
Regarding claim 3, Nakamura teaches wherein obtaining the first prompt further comprises: obtaining the first prompt based on the first image. [(see at least paragraph 36) “The communication terminal device 30 generates input information (for example, an image of an identification target object) and outputs the input information to the library presenting device 10. The communication terminal device 30 also presents to a user presentation content (for example, a library suitable for a defined task) output by the library presenting device 10. The communication terminal device 30 includes a communicator 31, a storage 32, a controller 33, an input information generator 34, and a user interface 35. Details of components of the communication terminal device 30 are described later.”]
Regarding claim 4, Nakamura teaches further comprising: selecting, in response to the first image indicating that the physical space comprises a plurality of second objects associated with the first object, the second object from the plurality of second objects. [(see at least paragraphs 48-50) “According to the present embodiment, the defined task is recognition of the target object and display of a recognition result. For example, when the defined task is executed with respect to an image including multiple target objects (such as industrial components), recognition (identification of bolts, nuts, springs, and the like) is executed with respect to each of the multiple target objects, and the recognition result (names of bolts, nuts, springs, and the like) is overlaid on the image and output.” As in 49 “The result of applying the defined task using the specified library is not limited to an actual execution result, and may be a result of calculation in a simulation. Simulation may include, for example, not only target object recognition, but also target object grasping and the like that is difficult to actually execute. For example, the result acquisition unit 132 may execute a 3D simulation according to the target object after recognizing the target object from the input information, and acquire a result of grasping. Further, a 3D model in the 3D simulation may be synthesized from, for example, CAD data, RGB information, depth information, and the like regarding the target object.”]
Regarding claim 5, Nakamura teaches further comprising: obtaining, in response to determining that the first image of the physical space indicates that the physical space comprises the first object but not the second object, the first object by the robotic device. [(see at least paragraph 48) “The result acquisition unit 132 acquires a result of applying a defined task to the target object included in the input information. In more detail, the result acquisition unit 132 acquires the result of the library generation management device 20 executing the defined task using a specified library. According to the present embodiment, the defined task is recognition of the target object and display of a recognition result. For example, when the defined task is executed with respect to an image including multiple target objects (such as industrial components), recognition (identification of bolts, nuts, springs, and the like) is executed with respect to each of the multiple target objects, and the recognition result (names of bolts, nuts, springs, and the like) is overlaid on the image and output.”]
Regarding claim 6, Nakamura teaches further comprising: providing, to the user, a message for querying the user for a location of the second object in the physical space; and obtaining, in response to receiving a response to the message from the user, the second object by the robotic device based on the response. [(see at least paragraph 50) “The evaluation acquisition unit 133 acquires an evaluation of the result of applying the defined task. According to the present embodiment, the evaluation is at least one of, but not limited to, precision or recall. The evaluation may include a visual judgment of “acceptable” or “unacceptable” from a user. The evaluation acquisition unit 133 may acquire the evaluation entered by a user into the communication terminal device 30. As another example, the evaluation acquisition unit 133 may acquire an expected result (at least one of expected precision or recall) entered by a user into the communication terminal device 30, rather than an evaluation itself. When acquiring the expected result for the defined task, the evaluation acquisition unit 133 may determine whether the result of applying the defined task based on the expected result is greater than or equal to the expected result and use that determination as the evaluation.”]
Regarding claim 12, Nakamura teaches a robotic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the robotic device to perform acts for performing a user task comprising: receiving, by the robotic device, a user task from a user, the user task instructing the robotic device to obtain a first object; determining a second object associated with the first object [(see at least paragraphs 37-39, 47-48) As in 37 “The robot control system 100 includes a robot 2 and a robot control device 110. The robot 2 moves a work target object 8 from a work start point 6 to a work target point 7 by means of an end effector 2B (grasping hand) at the end of an arm 2A. That is, the robot control device 110 controls the robot 2 so that the work target object 8 moves from the work start point 6 to the work target point 7. The robot control system 100 includes a camera 4. The camera 4 captures images of items and the like in an influence zone 5 that may affect operation of the robot 2.” As in 39 “the library group preferably includes many libraries so that many types of the work target object 8 may be recognized in the robot control system 100. However, when many libraries are present, selecting an appropriate library for the type of the work target object 8 may be difficult for a user. Further, before incorporating the robot control system 100, a user may want to know whether a target object can be properly recognized using a library in the library group. Further, even after the robot control system 100 has been incorporated, a user may want to know whether a new work target object 8 can be recognized.”]; and obtaining, in response to determining that a first image of a physical space in which the robotic device is located indicates that the physical space comprises the first object and the second object, the first object and the second object by the robotic device. [(see at least paragraph 37-40,48) As in 38 “The robot control device 110 recognizes the work target object 8 present in a space where the robot 2 works based on images captured by the camera 4. Recognition of the work target object 8 is executed using a library (trained model generated to recognize the work target object 8) included in the library group. The robot control device 110 downloads and deploys a library via the network 40 before executing recognition of the work target object 8.” As in 37 “FIG. 2 is a schematic diagram illustrating an example configuration of the robot control system 100. The robot control system 100 includes a robot 2 and a robot control device 110. The robot 2 moves a work target object 8 from a work start point 6 to a work target point 7 by means of an end effector 2B (grasping hand) at the end of an arm 2A. That is, the robot control device 110 controls the robot 2 so that the work target object 8 moves from the work start point 6 to the work target point 7. The robot control system 100 includes a camera 4. The camera 4 captures images of items and the like in an influence zone 5 that may affect operation of the robot 2.” As in 48 “According to the present embodiment, the defined task is recognition of the target object and display of a recognition result. For example, when the defined task is executed with respect to an image including multiple target objects (such as industrial components), recognition (identification of bolts, nuts, springs, and the like) is executed with respect to each of the multiple target objects, and the recognition result (names of bolts, nuts, springs, and the like) is overlaid on the image and output.”]
Regarding claim 15, This claim recites analogous limitations to claim 2 above, and is therefore rejected on the same premise.
Regarding claim 16, This claim recites analogous limitations to claim 3 above, and is therefore rejected on the same premise.
Regarding claim 17, This claim recites analogous limitations to claim 4 above, and is therefore rejected on the same premise.
Regarding claim 18, This claim recites analogous limitations to claim 5 above, and is therefore rejected on the same premise.
Regarding claim 19, This claim recites analogous limitations to claim 6 above, and is therefore rejected on the same premise.
Regarding claim 23, Nakamura teaches a non-transitory computer-readable storage medium having stored thereon a computer program which, when executed by a processor, causes the processor to implement acts for performing a user task comprising: receiving, by a robotic device, a user task from a user, the user task instructing the robotic device to obtain a first object; determining a second object associated with the first object [(see at least paragraphs 37-39, 47-48) As in 37 “The robot control system 100 includes a robot 2 and a robot control device 110. The robot 2 moves a work target object 8 from a work start point 6 to a work target point 7 by means of an end effector 2B (grasping hand) at the end of an arm 2A. That is, the robot control device 110 controls the robot 2 so that the work target object 8 moves from the work start point 6 to the work target point 7. The robot control system 100 includes a camera 4. The camera 4 captures images of items and the like in an influence zone 5 that may affect operation of the robot 2.” As in 39 “the library group preferably includes many libraries so that many types of the work target object 8 may be recognized in the robot control system 100. However, when many libraries are present, selecting an appropriate library for the type of the work target object 8 may be difficult for a user. Further, before incorporating the robot control system 100, a user may want to know whether a target object can be properly recognized using a library in the library group. Further, even after the robot control system 100 has been incorporated, a user may want to know whether a new work target object 8 can be recognized.”]; and obtaining, in response to determining that a first image of a physical space in which the robotic device is located indicates that the physical space comprises the first object and the second object, the first object and the second object by the robotic device. [(see at least paragraph 37-40,48) As in 38 “The robot control device 110 recognizes the work target object 8 present in a space where the robot 2 works based on images captured by the camera 4. Recognition of the work target object 8 is executed using a library (trained model generated to recognize the work target object 8) included in the library group. The robot control device 110 downloads and deploys a library via the network 40 before executing recognition of the work target object 8.” As in 37 “FIG. 2 is a schematic diagram illustrating an example configuration of the robot control system 100. The robot control system 100 includes a robot 2 and a robot control device 110. The robot 2 moves a work target object 8 from a work start point 6 to a work target point 7 by means of an end effector 2B (grasping hand) at the end of an arm 2A. That is, the robot control device 110 controls the robot 2 so that the work target object 8 moves from the work start point 6 to the work target point 7. The robot control system 100 includes a camera 4. The camera 4 captures images of items and the like in an influence zone 5 that may affect operation of the robot 2.” As in 48 “According to the present embodiment, the defined task is recognition of the target object and display of a recognition result. For example, when the defined task is executed with respect to an image including multiple target objects (such as industrial components), recognition (identification of bolts, nuts, springs, and the like) is executed with respect to each of the multiple target objects, and the recognition result (names of bolts, nuts, springs, and the like) is overlaid on the image and output.”]
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 7-10 and 20-22 are rejected under 35 U.S.C. 103 as being unpatentable over Nakamura in view of Ogawa (US 2020/0357382 A1).
Regarding claim 7, Nakamura has all of the elements of claim 1 as discussed above.
Nakamura does not explicitly teach wherein the first object is a book, and the method further comprises: receiving, from the user, a request to provide audio data associated with the book, and identifying first text data in the book by the robotic device; and converting the first text data into first audio data.
However, Ogawa teaches wherein the first object is a book, and the method further comprises: receiving, from the user, a request to provide audio data associated with the book, and identifying first text data in the book by the robotic device; and converting the first text data into first audio data. [(see at least paragraph 288) “ an oral computing device is provided, which includes a housing that holds at least: a memory device that stores thereon a data enablement application that includes multiple pairs of corresponding conversational bots and digital magazine modules, each pair specific to a user account and a topic; a display device to display a currently selected digital magazine; a microphone that is configured to record a user's spoken words as audio data; a processor configured to use the conversational bot to identify contextual data associated with the audio data, the contextual data including the currently selected digital magazine; a data communication device configured to transmit the audio data and the contextual data via a data network and, in response, receive response data, wherein the response data is text of an article about the topic; and an audio speaker that is controlled by the processor to read out the text of the article.”]
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of Nakamura to incorporate the teachings of Ogawa of wherein the first object is a book, and the method further comprises: receiving, from the user, a request to provide audio data associated with the book, and identifying first text data in the book by the robotic device; and converting the first text data into first audio data in order to process a variety of data and outputting digital media content, such as via audio. [(Ogawa 2)]
Regarding claim 8, Modified Nakamura has all of the elements of claim 7 as discussed above.
Nakamura does not explicitly teach further comprising: determining a position of the first text data in the book based on the request; and converting the first text data at the location to the first audio data.
However, Ogawa teaches further comprising: determining a position of the first text data in the book based on the request; and converting the first text data at the location to the first audio data. [(see at least paragraph 288) “an oral computing device is provided, which includes a housing that holds at least: a memory device that stores thereon a data enablement application that includes multiple pairs of corresponding conversational bots and digital magazine modules, each pair specific to a user account and a topic; a display device to display a currently selected digital magazine; a microphone that is configured to record a user's spoken words as audio data; a processor configured to use the conversational bot to identify contextual data associated with the audio data, the contextual data including the currently selected digital magazine; a data communication device configured to transmit the audio data and the contextual data via a data network and, in response, receive response data, wherein the response data is text of an article about the topic; and an audio speaker that is controlled by the processor to read out the text of the article.”]
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of modified Nakamura to further incorporate the teachings of Ogawa of determining a position of the first text data in the book based on the request; and converting the first text data at the location to the first audio data in order to output the full given text according to the audio style library to the user. [(Ogawa 286)]
Regarding claim 9, Modified Nakamura has all of the elements of claim 8 as discussed above.
Nakamura does not explicitly teach wherein the second object is a book, and the method further comprises: determining second text data associated with the first text data in the second object; and providing the second text data to the user.
However, Ogawa teaches wherein the second object is a book, and the method further comprises: determining second text data associated with the first text data in the second object; and providing the second text data to the user. [(see at least paragraph 295) “a memory device that stores thereon at least a data enablement application that includes multiple pairs of corresponding conversational bots and digital magazine modules, each pair specific to a user account and a topic; a display device to display a currently selected digital magazine; a microphone that is configured to record a user's spoken words as audio data; a processor configured to use the conversational bot to identify contextual data associated with the audio data, the contextual data including the currently selected digital magazine; a data communication device configured to transmit the audio data and the contextual data via a data network and, in response, receive response data, wherein the response data is text of an article about the topic; and an audio speaker that is controlled by the processor to output an audio response derived from at least the text of the article.”]
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of modified Nakamura to further incorporate the teachings of Ogawa of wherein the second object is a book, and the method further comprises: determining second text data associated with the first text data in the second object; and providing the second text data to the user in order to output the full given text according to the audio style library to the user. [(Ogawa 286)]
Regarding claim 10, Modified Nakamura has all of the elements of claim 9 as discussed above.
Nakamura does not explicitly teach further comprising: determining, in response to receiving a response to the second text data from the user, an evaluation of the response.
However, Ogawa teaches further comprising: determining, in response to receiving a response to the second text data from the user, an evaluation of the response. [(see at least paragraphs 301-302) As in 301 “In another example aspect, the audio response comprises a summarization of the text of the article and a question asking the user if they wish to hear the full article.” As in 302 “In another example aspect, the oral computing device further receives a yes response to the question, and the oral computing device then generates another audio response that includes reading out, via the audio speaker, the text of the article in entirety.”]
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the teachings of modified Nakamura to further incorporate the teachings of Ogawa of determining, in response to receiving a response to the second text data from the user, an evaluation of the response in order to understand questions and answers specific to various responses from the user. [(Ogawa 62)]
Regarding claim 20, This claim recites analogous limitations to claim 7 above, and is therefore rejected on the same premise.
Regarding claim 21, This claim recites analogous limitations to claim 8 above, and is therefore rejected on the same premise.
Regarding claim 22, This claim recites analogous limitations to claim 9 above, and is therefore rejected on the same premise.
The Examiner has cited particular paragraphs or columns and line numbers in the references applied to the claims above for the convenience of the Applicant. Although the specified citations are representative of the teachings of the art and are applied to specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested of the Applicant in preparing responses, to fully consider the references in their entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner. See MPEP 2141.02 [R-07.2015] VI. A prior art reference must be considered in its entirety, i.e., as a whole, including portions that would lead away from the claimed Invention. W.L. Gore & Associates, Inc. v. Garlock, Inc., 721 F.2d 1540, 220 USPQ 303 (Fed. Cir. 1983), cert, denied, 469 U.S. 851 (1984). See also MPEP §2123.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
(US 2021/0216767 A1) Yu - METHOD AND COMPUTING SYSTEM FOR OBJECT RECOGNITION OR OBJECT REGISTRATION BASED ON IMAGE CLASSIFICATION
(US 2024/0353853 A1) Jeong - TRANSPARENT OBJECT RECOGNITION AUTONOMOUS DRIVING SERVICE ROBOT MEANS
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MOHAMMED YOUSEF ABUELHAWA whose telephone number is (571)272-3219. The examiner can normally be reached Monday-Friday 8:30-5:00 with flex.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Wade Miles can be reached at 571-270-7777. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MOHAMMED YOUSEF ABUELHAWA/Examiner, Art Unit 3656
/WADE MILES/Supervisory Patent Examiner, Art Unit 3656