DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Remarks
35 USC § 101
Applicant’s remarks and amendments have been fully considered and have been determined to be unpersuasive, therefore the 35 USC § 101 is maintained.
Applicant argues that firstly, “Claim 1 as presently amended is not merely a "mental process" and is not directed to an abstract idea. In particular, claim 1 requires obtaining user input with gaze tracking sensors while the user is looking at a given media, determining an intent while the user is looking at the given media, obtaining a text summary using a trained model, and overlaying the text summary on a portion of the given media. Rather than an abstract idea, the limitations of claim 1 are directed to a specific technical improvement related to head-mounted devices (e.g., like in Enfish v. Microsoft, 822 F.3d 1327 (2016)) - determining a user's intent to view a given media using gaze tracking sensors, obtaining a text summary using a trained model, and overlaying the text summary on a portion of the given media is an unconventional method of determining text summaries in head-mounted devices.
However, the specific technical improvement in Enfish (namely the technical improvement to functioning of databases) is not analogous to the claimed technical improvement of the claimed invention because the claims are not directed to an improvement of the head mounted device itself and rather are directed to a method to be performed on a head mounted device which generic computer hardware configured to be worn on the head of a user.
Applicant further argues “claim 1 is tied to a specific structure (a head-mounted device with cameras, displays, processors, gaze tracking sensors), and displaying the text summary performs a specific function. In Data Engine Technologies LLC v. Google LLC, 906 F.3d 999 (Fed. Cir. 2018), the Federal Circuit held that a tab interface for a 3-D spreadsheet was not an abstract idea because the tab interface was a "specific structure (i.e., notebook tabs) within a particular spreadsheet display that performs a specific function (i.e., navigating within a three- dimensional spreadsheet)" (Id. at 1011). Here, claim 1 recites a specific structure (the head-mounted device and specific components in the head-mounted device) and a specific function (generating a text summary based on gaze tracking in the head- mounted device and overlaying the text summary on the summarized media using displays (s) in the head-mounted device). Like in Data Engine Technologies, the specific structure and function of claim 1 is not an abstract idea.”
However, the head mounted device as claimed in the amended Claim 1 is not analogous to the specific structure in Data Engine Technologies LLC v. Google LLC because in DataEngine the specific structure and specific function was found to represent an improvement to the functioning of computers and not because of the recitation of a specific structure and specific function. As highlighted above regarding Enfish, Claim 1 is not found to be directed to the improvement of the functioning of a head mounted device.
Applicant further argues “Claim 1 is not directed to an abstract idea that is a simple mental calculation/determination. In particular, using a trained model to obtain the text summary is not a "mental calculation" that can be performed in a user's mind at least because of the sheer amount of data that is required in a trained model. Therefore, obtaining the text summary using a trained model of claim 1 cannot be practically performed in the human mind and is therefore not an abstract mental process (see, e.g., RI Int'l, Inc. v. Cisco Systems, Inc., 930 F.3d 1295, 1304 (Fed. Cir. 2019)).”
However, the generation and obtaining of the text summary was the mental process at issue, not the use of a trained model. The trained model is a well understood routine and conventional additional element that is not significantly more than the mental process of summary generation. The human mind does not require a “sheer amount of data” to obtain a text summary.
Finally, applicant argues “Claim 1 recites an unconventional method of generating a text summary based on user intent and overlaying the text summary on the summarized media (e.g., claim 1 recites "significantly more" than an abstract idea), and is therefore patent eligible. In particular, the use of the gaze trackers to determine a user's intent which is then used in a trained model to obtain a text summary which in turn is overlaid on the summarized text is an unconventional way of summarizing text for a user of the head- mounted device. In other words, claim 1 involves "more than the performance of 'well-understood, routine, [and] conventional activities previously known to the industry'" (Aatrix Software19 v. Green Shades Software, 882 F.3d 1121 (2018)), and is therefore a patent-eligible inventive concept.”
However, claim 1 does not involve “more than the performance of 'well-understood, routine, [and] conventional activities previously known to the industry” as the only aspect considered under the well-understood, routine and conventional consideration of Step 2B are the additional elements and not the mental process of generating a summary. The additional elements, at issue, of gaze trackers and trained models are indeed well-understood, routine and conventional at the time of the effective filing date of the claimed invention.
35 USC § 103
Applicant argues “Terman fails to show or suggest a head- mounted device that uses gaze tracking sensors to obtain a user input that is used to determine the user's intent, as recited in claim 1 as presently amended. Instead, Terman discloses a user taking an image of text, which can then be summarized (see, e.g., paragraph 4). There is no suggestion in Terman that a user's intent is determined from a user input from a gaze tracker, much less that the summarized text is then overlaid on the original text, as recited in amended claim 1.”
Examiner agrees that Terman fails to disclose the head-mounted device, and intent from gaze tracker. However, Terman teaches summarized text is then overlaid on the original text (Summarize with summary on top of original document[0133]).
Specification
The disclosure is objected to because of the following informalities: Paragraph 00119 Lines 7-8
recites “FIG. 6”, however no Figure 6 is provided in the drawings.
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1,3-10,12-19,21-27 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claim(s) recite(s) elements which under their broadest reasonable interpretation are directed to mental processes. This judicial exception is not integrated into a practical application as explained below. The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception as explained below.
Regarding Claim 1, the claim recites a head-mounted device configured to be worn by a user, the head-mounted device comprising:
one or more cameras;
one or more displays;
one or more processors;
one or more gaze tracking sensors; and
memory storing instructions configured to be executed by the one or more processors, the instructions for:
with the one or more gaze tracking sensors, obtaining a user input while the user is looking at a given media in a physical environment;
determining an intent based on the user input while the user is looking at the given media; and in accordance with a determination that the intent represents an intent to provide a summary of text in the given media:
obtaining a text summary based on information regarding the text, the information regarding the text obtained from one or more images of the given media captured using the one or more cameras, wherein obtaining the text summary based on the information regarding the text comprises providing the information regarding the text to a trained model;
and presenting the text summary on the one more displays, wherein presenting the text summary comprises overlaying the text summary on a portion of the given media.
Claim Interpretation: Under the broadest reasonable interpretation, the terms of the claim are
presumed to have their plain meaning consistent with the specification as it would be interpreted by
one of ordinary skill in the art. See MPEP 2111.
Claim element b recites tracking a user’s gaze and obtaining a user input while they are looking at a given media. A human can track a user’s gaze and obtain an input from them when they are looking at a given media
Claim element c recites determining if an intent from the user if one regarding providing a summary of the text in the given media. A human can make such a determination
Claim element d recites obtaining a summary of the text in the given media. A human can summarize the text in a given media based on what they see
Claim element e recites overlaying the summary on a portion of the given media. A human can write down their summary using paper and pen and overlay the paper on the media in question.
Additional elements recited are: head-mounted device, camera, display, processor, gaze tracking sensor, memory, images, and trained model.
Claim 1 is patent illegible as it is directed to an abstract idea without being significantly more.
Step 1: This part of the eligibility analysis evaluates whether the claim falls within any statutory
category. See MPEP 2106.03. The claim is directed to a system, which falls within one of the statutory
categories of invention. (Step 1: YES).
Step 2A, Prong One: This part of the eligibility analysis evaluates whether the claim recites a
judicial exception. As explained in MPEP 2106.04, subsection II, a claim “recites” a judicial exception
when the judicial exception is “set forth” or “described” in the claim.
As discussed above, the broadest reasonable interpretation of elements (b)-(e) that those elements fall within the mental process groupings of abstract ideas because they cover concepts performed in the human mind, including observation, evaluation, judgment, and opinion . See MPEP 2106.04(a)(2), subsection III.
Claim element b is directed to a mental step as a human can track one’s gaze and receive an input. Claim element c is directed to a mental step as a human can make a determination if an intent corresponds to one requesting a summary. Claim element d is directed to a mental step as a human can summarize what they see. Claim element e is directed to a mental step as a human can write down the summary they created and overlay it on a media. Hence, these steps can be performed by a human, using “observation, evaluation, judgment, [and] opinion,” because they involve making determinations and identifications, which are mental tasks humans routinely do,' ” and thus can practically be performed in the human mind, In re Killian, 45 F.4th 1373, 1379 (Fed. Cir. 2022). Therefore, these limitations are considered together as an abstract idea for further analysis. (Step 2A, Prong One: YES).
Step 2A, Prong Two: This part of the eligibility analysis evaluates whether the claim as a whole
integrates the recited judicial exception into a practical application of the exception or whether the claim is “directed to” the judicial exception. This evaluation is performed by (1) identifying whether there are any additional elements recited in the claim beyond the judicial exception, and (2) evaluating those additional elements individually and in combination to determine whether the claim as a whole integrates the exception into a practical application. See MPEP 2106.04(d). The additional elements recited were head-mounted device, camera, display, processor, gaze tracking sensor, memory, images, and trained model.
These additional elements provide nothing more than mere instructions to implement an abstract idea on a generic computer. See MPEP 2106.05(f). MPEP 2106.05(f) provides the following considerations for determining whether a claim simply recites a judicial exception with the words “apply it” (or an equivalent), such as mere instructions to implement an abstract idea on a computer: (1) whether the claim recites only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished; (2) whether the claim invokes computers or other machinery merely as a tool to perform an existing process; and (3) the particularity or generality of the application of the judicial exception. The additional element of one or more images of the given media captured using the one or more cameras is an insignificant extra-solution activity towards writing a summary. Even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application (Step 2A, Prong Two: NO), and the claim is directed to the judicial exception. (Step 2A: YES).
Step 2B: This part of the eligibility analysis evaluates whether the claim as a whole amount to
significantly more than the recited exception, i.e., whether any additional element, or combination of
additional elements, adds an inventive concept to the claim. See MPEP 2106.05.
At Step 2A, the additional elements of head-mounted device, camera, display, processor, gaze tracking sensor, memory, images, and trained model were found to represent no more than mere instructions to apply the judicial exception to a computer using generic computer components. Mere
instructions to “apply” the abstract ideas, cannot provide an inventive concept. See MPEP 2106.05(f).
The analysis under Step 2A, Prong Two is carried through to Step 2B. Further, one or more images of
physical environment captured using the one or more cameras was found to be insignificant extra-
solution activity. However, a conclusion that an additional element is insignificant extra-solution activity
in Step 2A should be re-evaluated in Step 2B. See MPEP 2106.05, subsection I.A. At Step 2B, the re-
evaluation of the insignificant extra-solution activity consideration takes into account whether or not
the extra-solution activity is well understood, routine, and conventional in the field. See MPEP
2106.05(g). Taking a picture of text prior to obtaining information regarding the text is well
understand routine and conventional as disclosed in Ofgang’s publication on “Taking Notes vs.
Photographing Slides”. Therefore, this limitation remains insignificant extra solution activity even
upon reconsideration and does not amount to significantly more. Even when considered in
combination, these additional elements represent mere instructions to implement an abstract idea or
other exception on a computer and insignificant extra-solution activity, which do not provide an
inventive concept. (Step 2B: NO).
As such Claim 1 is patent illegible.
Regarding Claims 10 and 19, analysis analogous to that of Claim 1 is applicable.
Regarding Claim 3, 12, 21, a human can convert text they see into a format a machine can understand
Regarding Claims 4, 13, 22, the communication circuitry is an additional element which represent mere instructions to implement the abstract idea using generic computer components.
Regarding Claims 5, 14, 23, a human can create a summary by seeing an image of the given media.
Regarding Claims 6, 15, 24, a human can create a summary utilizing the age of the user in mind.
Regarding Claims 7, 16, 25, a human can adjust the length of a summary based on instructions from a user.
Regarding Claims 8, 17, 26, a human can select a subset of text for the basis of the summary they create
Regarding Claims 9, 18, 27, a human can make a selection and can position the overlayed text at a selected area.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 3-8, 10, 12-17, 19, 21-26 are rejected under 35 U.S.C. 103 as being unpatentable over Santoro(US PGPub 20220374645) in view of Terman (US PGPub 20130021346).
Regarding Claim 1, Santoro teaches a head-mounted device configured to be worn by a user(The embodiments disclosed herein may include or be implemented in conjunction with an artificial reality system… The artificial reality system that provides the artificial reality content may be implemented on various platforms, including a head-mounted display (HMD)[0141]), the head-mounted device comprising: one or more cameras(the on-board cameras[0008]); one or more displays(a head-mounted display (HMD)[0141]); one or more processors(Processor 1602 [Fig 16]); one or more gaze tracking sensors(eye tracking[0068], visual signals may comprise one or more gaze signals[0144]); and memory storing instructions configured to be executed by the one or more processors, the instructions for(Memory 1604 [Fig 16]):
with the one or more gaze tracking sensors(eye tracking[0068], visual signals may comprise one or more gaze signals[0144]), obtaining a user input(assistant system 140 may reactively perform contextual understanding of the visual inputs of text and objects after the user invokes the assistant system 140 for such assistance[0142]) while the user is looking at a given media in a physical environment(one or more visual signals may comprise one or more gaze signals indicating the first user is looking at the textual content[0144], images portraying textual content in a real-world environment associated with the first user[0143]);
determining an intent based on the user input while the user is looking at the given media(determining the user's potential intent based on the real-world text[0011]); and in accordance with a determination(The intent classifier 336b may determine the user's intent associated with the user input.[0098]) that the intent represents an intent to provide a summary of text in the given media(In particular embodiments, the assistant system may further assist the user to effectively and efficiently digest the obtained information by summarizing the information.[0062], The assistant system may…provide information based on the user input[0003]):
one or more images of the given media captured using the one or more cameras(images portraying textual content in a real-world environment associated with the first user[0143]), providing the information regarding the text to a trained model(The assistant system may then recognize, based on one or more machine-learning models and the one or more visual signals, the textual content.[0009]);
overlaying the text on a portion of the given media(the assistant system 140 may automatically recognize the text of the menu, translate it, and overlay AR effect of the same content in the user's native language[0158]).
Santoro does not teach obtaining a text summary based on information regarding the text, the information regarding the text obtained from one or more images of the given media captured using the one or more cameras, wherein obtaining the text summary based on the information regarding the text comprises providing the information regarding the text to a trained model; and presenting the text summary on the one more displays, wherein presenting the text summary comprises overlaying the text summary on a portion of the given media.
However, Terman teaches a determination that the user intent represents an intent to provide a
summary of text in a given media(with a single click of the camera provides a functional
summary of the textual images[Terman 0028]);, obtaining a text summary based on information regarding the text, the information regarding the text obtained from one or more images of the given media captured using the one or more cameras(Once the image is captured it is then processed seamlessly through OCR and text summarization programs resulting in a user understandable summary of the text in the captured images[Terman 0026]), wherein obtaining the text summary based on the information regarding the text comprises providing the information regarding the text to a trained model(The basis for Text Analyst's processing is a neural network technique that analyzes
preprocessed text to develop a semantic network. The semantic network provides the core
information required for clustering, summarizing, and navigating textual material [0090]); and presenting the text summary on the one more displays (the machine summary appears in the monitor screen of the device [Terman 0035]), wherein presenting the text summary comprises overlaying the text summary on a portion of the given media(Summarize with summary on top of original document[0133])..
It would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to combine the head-mounted device of Santoro with the summary generation of Terman because it would maximize "time on task" while minimizing the time-consuming steps of "concept capture" and "compression"(Terman Abstract).
Regarding Claim 3, Terman teaches wherein providing the information regarding the text to the
trained model comprises: performing optical character recognition (OCR) on the text(the software
automatically launches an OCR program which converts the image to editable text [0028] ), wherein
performing optical character recognition on the text comprises converting the text into machine-
readable text(the software automatically launches an OCR program which converts the image to
editable text [0028], HTML CSV TXT [0046]); and providing the machine-readable text to the trained
model(editable text which is then summarized by the machine summarizer software [0028] ).
Regarding Claim 4, Terman discloses communication circuitry, wherein providing the machine-
readable text to the trained model comprises providing the machine-readable text to the trained model
using the communication circuitry(Figure 1. document summarization as provided by SSW 111 may be
performed on any electronic documents that may be uploaded to the server from an internet -capable
appliance such as from computers 107a, 107b, and 107d [0156], Client station 103 is similar to station
101 in that a computer 107b running a browser instance 112b is provided as well as an OCR scanner
109b[0155]. network backbone 106. Backbone 106 represents all of the lines equipment and access
points making up the Internet network as a whole[0153] ).
Regarding Claim 5, Terman teaches providing the information regarding the text to the trained
model comprises providing the one or more images to the trained model(In some instances, the textual
material may be in a format such as WORD or pdf that is transferred directly and converted to a summary by the summarizer program without requiring the intermediary OCR processing [0040] ).
Regarding Claim 6, Terman discloses wherein the instructions further comprise instructions for:
providing contextual information to the trained model in addition to the information regarding the text, wherein the text summary is based on both the contextual information and the information regarding the text(The user can choose the format of the final summary[0077] ) and wherein the contextual information comprises location information, cultural information, reading level information, age information, subject matter knowledge information, education information, historical query information, user preference information(The user therefore first selects the summarizer intervals from a menu comprising: [0075]), temporal information, or calendar information.
Regarding Claim 7, Terman teaches wherein the instructions further comprise instructions for:
providing one or more response length parameters(A summary with a compression of 30% would be
expected to include more relevant concepts than one with a compression of 10% [0067]) to the trained model in addition to the information regarding the text, wherein the text summary is based on both the one or more response length parameters and the information regarding the text and wherein the one or more response length parameters comprises an absolute maximum for a length of the text summary, a relative maximum that is based on a length of the text(The user can compress or condense
the summary document to varying degrees determined as a percent of the total word volume of the
original material.[0067]), an absolute minimum for a length of the text summary, or a relative minimum
that is based on a length of the text.
Regarding Claim 8, Santoro in view of Terman teaches wherein providing information regarding the text to the trained model comprises providing information regarding only a subset of the text to the trained model(The program can automatically summarize individual subsections of the textual material containing relatively homogeneous subject matter [Terman 0034]), wherein the head-mounted device further comprises one or more input components, and wherein the instructions further comprise instructions for:
selecting the subset of the text based on input to the one or more input components(selection of UI elements via touch or gesture), or any other type of suitable user input that may be detected[Santoro 0006]), wherein the one or more input components comprises a gaze detection sensor(eye tracking[0068], visual signals may comprise one or more gaze signals[0144]), a microphone, or a button; and presenting an indicator on the one more displays that identifies the subset of the text(This is accomplished by using a highlighting the text to be summarized [Terman 0034]).
Claims 10, 12-17, 19, 21-26 are directed to a method and non-transitory computer readable medium which recite similar limitations to Claims 1, 3-8 above and are rejected similarly.
Claim(s) 9,18,27 are rejected under 35 U.S.C. 103 as being unpatentable over Santoro(US PGPub 20220374645) in view of Terman (US PGPub 20130021346) as applied in prior claims, in view of Bala(US PGPub 20120189190).
Regarding Claim 9, Santoro and Terman do not teach wherein the instructions further comprise
instructions for: selecting a position for the text summary on the one more displays based on the one or
more images of the physical environment, wherein presenting the text summary on the one more
displays comprises presenting the text summary on the one more displays at the selected position.
However, Bala teaches, wherein the instructions further comprise instructions for: selecting a
position for the text summary(receiving input regarding a user-selected location in the image for text insertion [0007]) on the one more displays based on the one or more images of the physical environment(receiving an image [0007]), wherein presenting the text summary on the one more displays comprises presenting the text summary on the one more displays at the selected
position(In another aspect, a computer implemented method for defining a bounding polygon in
an image for text insertion comprises receiving an image for personalization, receiving input regarding
a user-selected location in the image for text insertion [0007]).
It would have been obvious to a person having ordinary skill in the art before the effective filing
date of the claimed invention having the teachings of Santoro and Terman to further include the
placement of text at a selected position from Bala because the user can add value by personalizing
documents [Bala 0002]
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Djamasbi(US PGPub 20200159986) discloses utilizing an eye tracker and replacing the text in the user’s field of vision with simplified text(summary is simplified text). Additionally, Djamasbi utilizes personalized context info(age).
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ARJUN R SWAMY whose telephone number is (571)272-9763. The examiner can normally be reached Mon-Fri 8-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at (571) 272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ARJUN SWAMY/Examiner, Art Unit 2654
/HAI PHAN/Supervisory Patent Examiner, Art Unit 2654