DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
This action is in reply to the amendment filed on 5/19/2026.
Claims 1, 3, 8-9, and 11-12 have been amended and are hereby entered.
Claims 13-14 have been added.
Claims 2 and 4-7 have been canceled.
Claims 1, 3, and 8-14 are currently pending and have been examined.
This action is made FINAL.
Response to Applicant’s Arguments
Claim Interpretation
The amendments to Claim 1 obviate the need for the previous 112(f) interpretation of “control unit” therein; therefore, this interpretation is withdrawn.
Claim Rejections – 35 USC § 112
The present cancellations of Claims 6-7 obviate the 112(a) rejections thereto; therefore, these rejections are withdrawn.
The present cancellation of Claim 2 obviates the 112(b) rejections thereto; therefore, these rejections are withdrawn. The rolling up of features originally claimed in Claim 2 into the independent claims has been redrafted sufficiently to avoid similar 112(b) rejections.
Claim Rejections – 35 USC § 101
Applicant’s arguments regarding the 101 analysis have been considered and are unpersuasive.
Applicant argues that the independent claims as presently amended, particularly based upon the incorporation of modified subject matter of original Claim 4, embodies an improvement to a technology and is therefore eligible under Step 2A, Prong Two of the 101 subject matter eligibility analysis. More specifically, Applicant argues that the adding of image analysis results to the know-how information used by the AI/LLM to generate a result to a query of a user of the second terminal device constitutes “an improvement to the functioning of a computer, or an improvement to technical field of generative AI” as “it is possible to further reinforce the technical information (i.e., the know-how information), and it is highly likely to obtain a reply with high accuracy in terms of the generative AI.” Examiner disagrees, finding that what is described by Applicant does not constitute an improvement to a technology as of Applicant’s date of filing, but instead merely constitutes the application of pre-existing AI techniques to a particular field of use.
As a preliminary matter, Examiner notes that in the cited language above, generative AI is indeed a “technical field” as understood in the 101 subject matter eligibility analysis, but the know-how information both as claimed and as described in the original disclosure is not “technical information” within the meaning of “technical” as understood in the 101 subject matter eligibility analysis. While certain embodiments of such know-how (e.g., knowledge of steps for how to properly repair a particular piece of plant equipment) might colloquially be called “technical information,” this information is not “technical” (ie: representative of an additional element) as analyzed under 101 standards, instead representing an abstract, commercial concept under these standards. Additionally, Applicant’s arguments here contain reference to numerous features and details (e.g., in relation to Paragraphs 0099, 0101, and 0103) which are not captured by the claim language as presently drafted.
Even ignoring the uncertain language of the present Remarks and original disclosure (e.g., “may,” “it is possible to”) which appears to be an admission that the argued improvement is not necessarily achieved or at least is not always achieved, Applicant is essentially arguing that increasing the amount of input/training data provided to an AI/LLM application results in more accurate outputs of said application. This is a key trait of AI-based applications (including LLMs) since their inception, and indeed is a core function of such applications. This is not an improvement of Applicant’s invention here, but rather a long pre-existing feature of AI technology writ large. Merely specifying the particular field of use to which the present invention applies AI technology (ie: simulating the knowledge of SMEs to provide responses to queries as such SMEs might) and the form of this input/training data (ie: know-how information, including supplementing such know-how information with the results of an image analysis of an image contained within such know-how information) does not make this otherwise. As such, one of ordinary skill in the art could not reasonably conclude that this constitutes an improvement to a technology as of Applicant’s effective filing date.
Claim Rejections – 35 USC § 102/103
Applicant’s arguments regarding the 102 and 103 analyses have been considered and are unpersuasive.
The independent claims are presently amended to incorporate modified subject matter of original Claims 2, 4, and 5, and all present art-based arguments are based on these amendments.
Regarding the previous 102 rejections, Applicant is correct that the incorporation of the subject matter of Claim 5, which was previously rejected under 103 standards, obviates the previous 102 rejections of the independent claims.
Regarding 103 standards, Applicant’s first argument is based on the presently rolled up subject matter of original Claim 4, particularly arguing that the primary Izquierdo-Domenech reference does not properly read upon the limitation “perform image analysis related to the first image that is included in the know-how information.” To this end, Applicant argues that “[t]he ray-casting mentioned in Izquierdo-Domenech calculates a line or ray from 2D touch coordinates with the direction derived from the camera frustrum (see page 4, left column, last paragraph of Izquierdo-Domenech), in order to select the object closest to the user” and “[m]oreover, this feature relates to the 3D scanned mesh from the virtual environment.” Examiner disagrees, noting that this explanation of content of Izquierdo-Domenech is at best reductive and incomplete.
Applicant’s explanation here ignores that the “2D touch coordinates” relating to a “3D scanned mesh from the virtual environment” in fact relates to image analysis of an image captured on a device of the SME, wherein the ray-casting functionality is used to identify a particular area within that image of the SME’s surroundings to which SME intends entered notes or information to relate. While Applicant is not entirely incorrect that this “relates to the 3D scanned mesh from the virtual environment,” this virtual environment is directly correlated to the shop floor in which the SME and the shop floor operator are working, with the particular location or piece of equipment identified using this ray-tracing functionality being specifically associated with the “pill” of knowledge being created. In other words, this functionality analyzes an image comprising this “pill,” and appends this “pill” with particular location information captured in said image such that it may be retrieved and used based on future images/queries of the same location or equipment identified by said ray-casting functionality. This falls well within the claim language of “perform image analysis related to the first image that is included in the know-how information” and “add a result of the image analysis to the know-how information.” While it is not entirely clear to Examiner from these Remarks why and how Applicant believes this disclosure does not read upon the argued claim language, Examiner’s best guess is that Applicant fails to properly consider the full scope of what is captured by this fairly vague and high-level claim language in view of the broadest reasonable interpretation standard.
Next, Applicant argues against the disclosure of functionality of original Claims 4-5, presently amended in a modified and combined fashion into the independent claims as “add a result of the image analysis to the know-how information, the result of the image analysis including information related to a position of a tool that is used by the person engaged in the work at the time of the work, a posture of the tool, and a point of application of the tool.”
Applicant’s first assertion to this effect is that the imaging and 3D modeling of tools and the usage/context thereof in the Jo reference, in isolation, “is not part of the know-how information.” In so doing, Applicant seemingly demands some at least limited application of 102 standards, ignoring that Claim 5 was previously rejected under 103 standards and in particular the content of original Claim 5 was previously rejected as a combination of Izquierdo-Domenech and Jo, not as being disclosed solely by Jo. While Examiner disputes Applicant’s contention that this information in Jo would not constitute “know-how information” given its usage in providing guidance and training to workers in performing equipment maintenance, this is irrelevant as Jo was not previously and is not presently cited as disclosing the “know-how information” piece of this claim language. Instead, Izquierdo-Domenech was and is cited against this feature. Applicant completely ignores how the content of Jo is used to modify the base disclosure of Izquierdo-Domenech in the previous 103 rejections, instead essentially arguing that the Jo reference does not teach something for which it was never cited. One cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references (see In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986)).
To the limited extent the actual standards of 103 are considered in this argument, Applicant merely asserts that “[e]ven if the skilled person would have combined the teaching of Jo with the teaching of Izquierdo-Domenech, the skilled person would not have in an obvious manner obtained a work support apparatus as defined in claims 1, 11, and 12 as amended,” this contention entirely based on the above-refuted arguments regarding the image analysis functionality of original Claim 4. As the basis of this argument is inaccurate as already explained above (ie: Izquierdo-Domenech does indeed contain disclosure reading upon the analysis of an image of the know-how information, and the adding of the results of such image analysis to the know-how information), the reference to this purported shortcoming of Izquierdo-Domenech does nothing to support Applicant’s contention here. Here again, Applicant ignores the actual content of the previous 103 rejections, including how Jo is used to modify the teachings of Izquierdo-Domenech, instead arguing that each of Izquierdo-Domenech and Jo in isolation do not include particular details of the claim language to which they were never applied.
The remainder of Applicant’s arguments here constitute bare conclusory statements that the cited art does not disclose particular presently amended language, providing no explanation of Applicant’s supposed disagreement with how the disclosure of such references differ from said amended language. This constitutes an improper argument, giving Examiner no understanding of Applicant’s thought process and nothing particular to which Examiner may respond. Examiner notes generally that he disagrees with these unexplained conclusory statements, finding that the Izquierdo-Domenech reference discloses image analysis results in the manner explained above, said results being used in the generation of the reply to a query of the shop floor operator.
Claim Objections
Claim 14 is objected to because of the following informality: “…the presentation information that is present to the second terminal device…” should read “…the presentation information that is presented to the second terminal device…” Appropriate correction is required.
Claim Rejections – 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1, 3, and 8-14 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding Claims 1, 11, and 12, the limitations of managing a work support operation of work related to at least a plant or infrastructure; acquire know-how information on work related to at least a plant or infrastructure from the work performed by a person engaged in the work in a work site, wherein the know-how information includes at least a first image related to the work, a first speech related to the work, and a first text that has been converted from the first speech, all of which have been recorded by the first terminal device; accumulate the know-how information; generate, when inquiry information related to the work is accepted, based on the inquiry information, a question that requests generation of a reply produced using the know-how information, the inquiry information includes at least a second image related to the work, a second speech related to the work, and a second text that has been converted from the second speech, all of which have been recorded by the second terminal device; generate presentation information that is presented to a second terminal device corresponding to an inquiry source based on the reply to the question received from the generative AI; transmit the presentation information to the second terminal device; transmits the know-how information to the server device in a case where the first terminal device is in a first mode; transmits the inquiry information to the server device in a case where the second terminal device is in a second mode; receives the presentation information with respect to the inquiry information from the server device; presents the presentation information to another person engaged in the work; add a result of the image analysis to the know-how information, the result of the image analysis including information related to a position of a tool that is used by the person engaged in the work at the time of the work, a posture of the tool, and a point of application of the tool; and generate the question represented by at least one of the second image, the second speech, and the second text according to a modality capable of being accepted by the large language model, such that the reply is generated using the know-how information to which the result of the image analysis is added, as drafted, are processes that, under their broadest reasonable interpretations, cover certain methods of organizing human activity. For example, these limitations fall at least within the enumerated categories of commercial or legal interactions and/or managing personal behavior or relationships or interactions between people (see MPEP 2106.04(a)(2)(II)).
Additionally, the limitations of managing a work support operation of work related to at least a plant or infrastructure; acquire know-how information on work related to at least a plant or infrastructure from the work performed by a person engaged in the work in a work site, wherein the know-how information includes at least a first image related to the work, a first speech related to the work, and a first text that has been converted from the first speech, all of which have been recorded by the first terminal device; accumulate the know-how information; generate, when inquiry information related to the work is accepted, based on the inquiry information, a question that requests generation of a reply produced using the know-how information, the inquiry information includes at least a second image related to the work, a second speech related to the work, and a second text that has been converted from the second speech, all of which have been recorded by the second terminal device; generate presentation information that is presented to a second terminal device corresponding to an inquiry source based on the reply to the question received from the generative AI; transmit the presentation information to the second terminal device; transmits the know-how information to the server device in a case where the first terminal device is in a first mode; transmits the inquiry information to the server device in a case where the second terminal device is in a second mode; receives the presentation information with respect to the inquiry information from the server device; presents the presentation information to another person engaged in the work; add a result of the image analysis to the know-how information, the result of the image analysis including information related to a position of a tool that is used by the person engaged in the work at the time of the work, a posture of the tool, and a point of application of the tool; and generate the question represented by at least one of the second image, the second speech, and the second text according to a modality capable of being accepted by the large language model, such that the reply is generated using the know-how information to which the result of the image analysis is added, as drafted, are processes that, under their broadest reasonable interpretations, cover mental processes. For example, these limitations recite activity comprising observations, evaluations, judgments, and opinions (see MPEP 2106.04(a)(2)(III)).
If a claim limitation, under its broadest reasonable interpretation, covers fundamental economic principles or practices, commercial or legal interactions, managing personal behavior or relationships, or managing interactions between people, it falls within the “Certain Methods of Organizing Human Activity” grouping of abstract ideas. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind or with the aid of pen and paper but for recitation of generic computer components, it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea.
The judicial exception is not integrated into a practical application. In particular, the claim recites the additional elements of a work support apparatus/a server device/a computer comprising a processor and a storage device, the processor being configured to perform functions and the storage device configured to store information; a first terminal device; generative AI that uses a large language model, said large language model capable of accepting inputs of one or more modalities; a second terminal device; and perform image analysis related to the first image that is included in the know-how information. A work support apparatus/a server device/a computer comprising a processor and a storage device, the processor being configured to perform functions and the storage device configured to store information; a first terminal device; generative AI that uses a large language model, said large language model capable of accepting inputs of one or more modalities; and a second terminal device, in the context of the claims as a whole, amount to no more than mere instructions to apply a judicial exception (see MPEP 2106.05(f)). Perform image analysis related to the first image that is included in the know-how information, in the context of the claims as a whole, amounts to no more than insignificant extra-solution activity (see MPEP 2106.05(g)). Accordingly, these additional elements do not integrate the abstract ideas into a practical application because they do not, individually or in combination, impose any meaningful limits on practicing the abstract ideas. The claims are therefore directed to an abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the judicial exception into a practical application, the additional elements amount to no more than mere instructions to apply a judicial exception, and insignificant extra-solution activity for the same reasons as discussed above in relation to integration into a practical application. The limitation found to recite insignificant extra-solution activity is further found to be well-understood, routine, and conventional as per the standards of 112(a) as one of ordinary skill in the art at the time of filing would recognize it as such in view of the high-level description of this functionality in the original disclosure (see, e.g., Paragraphs 0099, 0103, and 0114 as filed). These cannot provide an inventive concept. Therefore, when considering the additional elements alone and in combination, there is no inventive concept in the claims, and thus the claims are not patent eligible.
Claims 3, 8-10, and 13-14, describing various additional limitations to the device of Claim 1, amount to substantially the same unintegrated abstract idea as Claim 1 (upon which these claims depend, directly or indirectly) and are rejected for substantially the same reasons.
Claim 3 additionally discloses wherein the large language model is a multimodal large language model (generally linking the use of a judicial exception to a particular technological environment or field of use), which does not integrate the claim into a practical application.
Claim 8 additionally discloses wherein the first terminal device and the second terminal device are augmented reality (AR) terminals (generally linking the use of a judicial exception to a particular technological environment or field of use), which does not integrate the claim into a practical application.
Claim 9 additionally discloses wherein the processor is configured to generate the presentation information such that the presentation information is superimposed onto a real space in an AR display space of the AR terminal (insignificant extra-solution activity), which does not integrate the claim into a practical application. The limitation found to recite insignificant extra-solution activity is further found to be well-understood, routine, and conventional as per the standards of 112(a) as one of ordinary skill in the art at the time of filing would recognize it as such in view of the high-level description of this functionality in the original disclosure (see, e.g., Paragraphs 0021-0022, 0080, and 0118-0119 as filed).
Claim 10 additionally discloses wherein the know-how information includes technical information that is related to the work and that is provided by the person engaged in the work who is a skilled technical person (further defining the abstract idea already set forth in Claim 1), which does not integrate the claim into a practical application.
Claim 13 additionally discloses wherein the processor is further configured to associate the first image, the first speech and the first text as a data group (an abstract idea in the form of a certain method of organizing human activity and a mental process); and assign various kinds of tag data to the data group (an abstract idea in the form of a certain method of organizing human activity and a mental process), which do not integrate the claim into a practical application.
Claim 14 additionally discloses wherein the storage device stores presentation information generation information that includes various kinds of parameters that are used when the presentation information that is present to the second terminal device is generated (an abstract idea in the form of a certain method of organizing human activity and a mental process); and the presentation information generation information includes information related to a coordinate system of an AR display space included in a display of the second terminal device (generally linking the use of a judicial exception to a particular technological environment or field of use), which do not integrate the claim into a practical application.
Claim Rejections – 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1 and 8-14 are rejected under 35 U.S.C. 103 as being unpatentable over Izquierdo-Domenech et al, "Large Language Models for in Situ Knowledge Documentation and Access With Augmented Reality," Int'l Journal of Interactive Multimedia and Artificial Intelligence, Vol. 9, No. 3 (hereafter, “Izquierdo-Domenech”) in view of Jo et al (KR 102104326) (hereafter, “Jo”).
Regarding Claims 1, 11, and 12, Izquierdo-Domenech discloses:
a work support apparatus/a server device/a computer comprising a processor and a storage device, the processor being configured to perform functions and the storage device configured to store information (Abstract; pg. 4; Appendix; support system for shop floor operations in industrial settings; this appendix illustrates different requests the application can make to the server; AR systems can reduce the cost of having SMEs on-site; one of ordinary skill in the art would readily recognize such a server/computer as comprising at least a processor and a memory);
a first terminal device that is used for the work performed by a person engaged in the work in a work site (pgs. 4, 6; Figs. 1 and 3; work-related information on an industrial shop floor is acquired from an AR device or mobile device used by the subject matter expert (SME); mobile application developed for use with both Android and iOS devices; AR devices such as depicted in Figs. 2-3; Two shop floor roles: SMEs add information, operator retrieves it via NL/AR);
a second terminal device (pgs. 4, 6; Figs. 2-3; once the site has been enriched with anchored "pills" of knowledge, it is time for the shop floor operator to utilize these resources, reducing the frequency of their need to seek clarification from the SME and technical documentation; Speech Recognition for NL Queries; mobile application developed for use with both Android and iOS devices; AR devices such as depicted in Figs. 2-3; Two shop floor roles: SMEs add information, operator retrieves it via NL/AR);
acquire know-how information on work related to at least a plant or infrastructure from a first terminal device that is used for the work performed by a person engaged in the work in a work site (pgs. 4, 6; Fig. 1; Section A. SME: Context Enrichment With Information; work-related information on an industrial shop floor is acquired from an AR device or mobile device used by the subject matter expert (SME); mobile application developed for use with both Android and iOS devices; AR devices such as depicted in Figs. 1-3);
accumulate the know-how information in the storage device (pgs. 4, 6; Appendix; work-related information on an industrial shop floor is acquired from an AR device or mobile device used by the subject matter expert (SME); persisting of the information on the server-side using a Python FastAPI framework; information is added to repository, and checked against previously recorded information for consistency and/or redundancy; Create new note; Store begin/end indices for note in transformer context);
generate, when the work support apparatus accepts inquiry information related to the work, based on the inquiry information, a question that requests generative AI that uses a large language model to generate a reply produced using the know-how information (Abstract; pgs. 2, 4; Figs. 2-3, 5; Appendix; our method employs Large Language Models (LLMs) to allow experts to describe elements from the real environment in NL and select corresponding AR elements in a dynamic and iterative process; an ideal solution would offer multiple interaction options, including operators' ability to ask questions and receive answers in Natural Language (NL), as discussed in [12]; this information "pills" will be used by the system with two purposes: 1. To retrieve a specific "pill" linked to a specific person in the environment as-is, and 2. For obtaining answers to specific questions; paraphrasing the input questions; using architectures that have already been trained with vast amounts of data allows the pre-trained transformers to have already learned most of the semantics of NL, so they can process and answer most of the questions or suggestions that the user asks in NL; this appendix illustrates different requests the application can make to the server; uses of a GPT-JT model to generate a reply based on the know-how information of the SME);
generate presentation information that is presented to a second terminal device corresponding to an inquiry source based on the reply to the question received from the generative AI (pg. 4; Figs. 2-3, 5; in response to region selection/speech recognition & question in natural language, the system generates a response, e.g., incorporating a notes list or answers to received questions, said answers anchored to physical locations; presentation of answer to the question on the terminal device of the operator, potentially in the form of AR feedback);
transmit the presentation information to the second terminal device (pg. 4; Figs. 2-3, 5; in response to region selection/speech recognition & question in natural language, the system generates a response, e.g., incorporating a notes list or answers to received questions, said answers anchored to physical locations; presentation of answer to the question on the terminal device of the operator, potentially in the form of AR feedback);
wherein the know-how information includes at least a first image related to the work, a first speech related to the work, and a first text that has been converted from the first speech, all of which have been recorded by the first terminal device (pgs. 4, 6; Figs. 1-3; Sections A. SME: Context Enrichment With Information and D: Consistency of the Information; know-how input using AR images, speech and text converted from the speech; Speech Recognition using Vosk toolkit; speech-to-text conversion);
the processor is further configured to: perform image analysis related to the first image that is included in the know-how information (pgs. 2-4; the "pills" of SME knowledge are anchored in the physical space; leverage the knowledge and expertise of SMEs to create dynamic environments, enhancing them with knowledge anchors into spatial 3D real environments to improve efficiency and profitability; Section A. SME: Context Enrichment With Information; mapping of the pills of information to the environment using ray-casting techniques);
the inquiry information includes at least a second image related to the work, a second speech related to the work, and a second text that has been converted from the second speech, all of which have been recorded by the second terminal device (pgs. 4, 6; Figs. 2-3; Section B. Shop Floor Operator: Information Retrieval; querying using AR images, speech and text converted from the speech; Speech Recognition for NL Queries); and
the processor is further configured to generate the question represented by at least one of the second image, the second speech, and the second text according to a modality capable of being accepted by the large language model, such that the generative AI generates the reply using the know-how information to which the result of the image analysis is added (pgs. 2-4, 6; Figs. 2-3; the "pills" of SME knowledge are anchored in the physical space; leverage the knowledge and expertise of SMEs to create dynamic environments, enhancing them with knowledge anchors into spatial 3D real environments to improve efficiency and profitability; Sections A. SME: Context Enrichment With Information, and B. Shop Floor Operator: Information Retrieval; mapping of the pills of information to the environment using ray-casting techniques; know-how input and querying using AR images, speech and text converted from the speech; paraphrasing the input questions; using architectures that have already been trained with vast amounts of data allows the pre-trained transformers to have already learned most of the semantics of NL, so they can process and answer most of the questions or suggestions that the user asks in NL).
Izquierdo-Domenech additionally discloses add a result of the image analysis to the know-how information (pgs. 2-4; the "pills" of SME knowledge are anchored in the physical space; leverage the knowledge and expertise of SMEs to create dynamic environments, enhancing them with knowledge anchors into spatial 3D real environments to improve efficiency and profitability; Section A. SME: Context Enrichment With Information; mapping of the pills of information to the environment using ray-casting techniques). Izquierdo-Domenech does not explicitly disclose but Jo does disclose wherein the result of the image analysis includes a position of a tool that is used by the person engaged in the work at the time of the work, a posture of the tool, and a point of application of the tool (Abstract; ¶ 0010-0014, 0041-0046, 0067-0069; an augmented reality based maintenance training system and method thereof; a camera unit for shooting a live-action image; a control unit for checking whether the tool model is used according to the maintenance procedure based on the identification information received from the tool model, and confirming whether an accurate contact is made to the maintenance position; a control unit that transmits its own identification information to the augmented reality device when the use is detected by the user through the sensor unit, and transmits the attitude and orientation detection values of the tool model detected through the sensor unit to the augmented reality device in real time).
It would have been obvious to one of ordinary skill in the art before the filing date of the claimed invention to include the maintenance training-related image analysis techniques of Jo with the AI- and LLM-based expertise data collection and question answering system of Izquierdo-Domenech because the combination merely applies a known technique to a known device/method/product ready for improvement to yield predictable results (see KSR Int’l Co. v. Teleflex, Inc., 550 U.S. 398, 415-421 (2007) and MPEP 2143). The known techniques of Jo are applicable to the base device (Izquierdo-Domenech), the technical ability existed to improve the base device in the same way, and the results of the combination are predictable because the function of each piece (as well as the problems in the art which they address) are unchanged when combined.
Regarding Claim 8, Izquierdo-Domenech in view of Jo discloses the limitations of Claim 1. Izquierdo-Domenech additionally discloses wherein the first terminal device and the second terminal device are augmented reality (AR) terminals (Title; Abstract; pg. 4; Figs. 2-3; Augmented reality (AR) has become a powerful tool for assisting operators in complex environments, such as shop floors, laboratories, and industrial settings; the presented architecture implementation relies on the fact that the environment needs to be previously scanned, a common feature in current SLAM-based AR solutions; using ray-casting techniques, alongside touch interaction in AR, enables SMEs to pinpoint and enrich specific features of the 3D scanned mesh from the virtual environment; system presents AR feedback to received questions).
Regarding Claim 9, Izquierdo-Domenech in view of Jo discloses the limitations of Claim 8. Izquierdo-Domenech additionally discloses wherein the processor is configured to generate the presentation information such that the presentation information is superimposed onto a real space in an AR display space of the AR terminal (Title; Abstract; pg. 4; Figs. 2-3; system presents AR feedback to received questions, e.g., "touchable" anchor points for selection/linked to answer).
Regarding Claim 10, Izquierdo-Domenech in view of Jo discloses the limitations of Claim 1. Izquierdo-Domenech additionally discloses wherein the know-how information includes technical information that is related to the work and that is provided by the person engaged in the work who is a skilled technical person (pg. 4; Fig. 1; Section A. SME: Context Enrichment With Information; work-related information on an industrial shop floor is acquired from an AR device or mobile device used by the subject matter expert (SME); using ray-casting techniques, alongside touch interaction in AR, enables SMEs to pinpoint and enrich specific features of the 3D scanned mesh from the virtual environment).
Regarding Claim 13, Izquierdo-Domenech in view of Jo discloses the limitations of Claim 1. Izquierdo-Domenech additionally discloses:
wherein the processor is further configured to associate the first image, the first speech and the first text as a data group (pgs. 2-4, 6; Figs. 1-3; Sections A. SME: Context Enrichment With Information and D: Consistency of the Information; know-how input using AR images, speech and text converted from the speech; Speech Recognition using Vosk toolkit; speech-to-text conversion; "pills" of SME knowledge); and
assign various kinds of tag data to the data group (pgs. 2-4; Appendix; the "pills" of SME knowledge are anchored in the physical space; leverage the knowledge and expertise of SMEs to create dynamic environments, enhancing them with knowledge anchors into spatial 3D real environments to improve efficiency and profitability; SME adds notes to the environment; non-structured information anchored to specific elements; Section A. SME: Context Enrichment With Information; mapping of the pills of information to the environment using ray-casting techniques; Create a new note; Store begin/end indices for note in transformer context; Update model; TABLE I. GPT-JT Tests Label Information as "New", "Redundant", or "Contradictory" Based on Context).
Regarding Claim 14, Izquierdo-Domenech in view of Jo discloses the limitations of Claim 1. Izquierdo-Domenech additionally discloses:
wherein the storage device stores presentation information generation information that includes various kinds of parameters that are used when the presentation information that is present to the second terminal device is generated (pgs. 2-4, 6; Figs. 1-3; Appendix; the "pills" of SME knowledge are anchored in the physical space; leverage the knowledge and expertise of SMEs to create dynamic environments, enhancing them with knowledge anchors into spatial 3D real environments to improve efficiency and profitability; SME adds notes to the environment; non-structured information anchored to specific elements; Section A. SME: Context Enrichment With Information; mapping of the pills of information to the environment using ray-casting techniques; Create a new note; Store begin/end indices for note in transformer context; once the site has been enriched with anchored "pills" of knowledge, it is time for the shop floor operator to utilize these resources, reducing the frequency of their need to seek clarification from the SME and technical documentation; the retrieval of these “pills” and the contents thereof in response to a shop floor operator query is performed based on this spatial anchoring); and
the presentation information generation information includes information related to a coordinate system of an AR display space included in a display of the second terminal device ("pgs. 2-4, 6; Figs. 1-3; Appendix; the ""pills"" of SME knowledge are anchored in the physical space; leverage the knowledge and expertise of SMEs to create dynamic environments, enhancing them with knowledge anchors into spatial 3D real environments to improve efficiency and profitability; SME adds notes to the environment; non-structured information anchored to specific elements; Section A. SME: Context Enrichment With Information; mapping of the pills of information to the environment using ray-casting techniques; once the site has been enriched with anchored ""pills"" of knowledge, it is time for the shop floor operator to utilize these resources, reducing the frequency of their need to seek clarification from the SME and technical documentation; since many anchors might be disseminated around the site, the shop floor operator has the option to select a rectangular area and retrieve all the ""pills"" within that selection; applying the ray-casting methods as introduced in section III.A, the system can obtain the 3D virtual coordinates of a chosen location on the shop floor by utilizing touch interaction; the process involves the shop floor drawing an area of interest from which anchored ""pills"" can be retrieved").
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Izquierdo-Domenech in view of Jo and Yongquiang et al, "Enhancing the Spatial Awareness Capability of Multi-Modal Large Language Model," arXiv:2310.20357, Cornell Univ. Library (hereafter, “Yongquiang”).
Regarding Claim 3, Izquierdo-Domenech in view of Jo discloses the limitations of Claim 1. Izquierdo-Domenech does not explicitly disclose but Yongquiang does disclose wherein the large language model is a multimodal large language model (Abstract; Introduction; Multi-Modal Large Language Model (MLLM);).
The rationale to combine Izquierdo-Domenech and Jo remains the same as for Claim 1. It would have been obvious to one of ordinary skill in the art before the filing date of the claimed invention to include the multi-modal LLM-based structure and question answering techniques of Yongquiang with the AI- and LLM-based expertise data collection and question answering system of Izquierdo-Domenech because the combination merely applies a known technique to a known device/method/product ready for improvement to yield predictable results (see KSR Int’l Co. v. Teleflex, Inc., 550 U.S. 398, 415-421 (2007) and MPEP 2143). The known techniques of Yongquiang are applicable to the base device (Izquierdo-Domenech), the technical ability existed to improve the base device in the same way, and the results of the combination are predictable because the function of each piece (as well as the problems in the art which they address) are unchanged when combined.
Discussion of Prior Art Cited but Not Applied
For additional information on the state of the art regarding the claims of the present application, please see the following documents not applied in this Office Action (all of which are prior art to the present application):
PGPub 20250021919, claiming the benefit of Provisional 63526947 – “Enterprise Knowledge Retention and Access System,” Belkin et al, disclosing a system for gathering work information of industry experts, using said information to create generative AI-based digital twins of such experts, and using said digital twins to answer questions in place of such experts/simulate the expertise of such experts
PGPub 20240419950 – “Systems, Devices, and Methods for Enterprise System Integration Using Machine Learning,” Tong et al, disclosing a system for building a plurality of AI-based subject experts based on data sources, and filter a request to a corresponding AI expert to answer said request
PGPub 20230274095 – “Autonomous Conversational AI System Without Any Configuration by a Human,” Kelkar et al, disclosing a system for building generative AI-based topic-specific conversation models, and using such models to answer user questions
Kovalenko et al, Opportunities and Challenges to Integrate Artificial Intelligence Into Manufacturing Systems: Thoughts From a Panel Discussion, IEEE Robotics & Automation Magazine, Vol. 30, Issue 2, disclosing techniques for integrating AI systems into manufacturing environments, including to answer questions using the knowledge of subject matter experts
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MARK C CLARE whose telephone number is (571)272-8748. The examiner can normally be reached Monday-Friday 6:30am-2:30pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jeffrey Zimmerman can be reached at (571) 272-4602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MARK C CLARE/Examiner, Art Unit 3628
/JEFF ZIMMERMAN/Supervisory Patent Examiner, Art Unit 3628