Prosecution Insights
Last updated: October 04, 2026
Application No. 18/442,910

GAZE BASED DICTATION

Final Rejection §101§103
Filed
Feb 15, 2024
Priority
Sep 03, 2021 — provisional 63/240,696 +3 more
Examiner
OGUNBIYI, OLUWADAMILOL M
Art Unit
2653
Tech Center
2600 — Communications
Assignee
Apple Inc.
OA Round
4 (Final)
77%
Grant Probability
Favorable
5-6
OA Rounds
3m
Est. Remaining
96%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
243 granted / 315 resolved
+15.1% vs TC avg
Strong +19% interview lift
Without
With
+19.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
28 currently pending
Career history
342
Total Applications
across all art units

Statute-Specific Performance

§101
20.8%
-19.2% vs TC avg
§103
49.9%
+9.9% vs TC avg
§102
11.2%
-28.8% vs TC avg
§112
13.3%
-26.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 315 resolved cases

Office Action

§101 §103
DETAILED ACTION Claims 1, 3 – 9 and 11 – 37 are pending. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment With regard to the Non-Final Office Action from 25 March 2026, the Applicant has filed a response on 22 June 2026. New claims — 22 – 37 are added. Response to Arguments With regard to the 35 U.S.C. 101 rejection given to the claims for being directed to a judicial exception without significantly more, the Applicant argues (Remarks: page 12 par 6 – page 13 par 3) that the [independent] claims as currently presented, do not recite a mental process, as ‘the human mind is not trained to perform the specific task of entering an editing mode based on specific criteria including “determining that the utterance includes one or more predetermined words and that a velocity of a detected gaze of a user over a word displayed on a screen of a device is below a predetermined threshold”’ stating further that a person would consider the information as a whole without the individual structures mentioned in the claims. The Examiner agrees that a human would consider the entire information as a whole while deciding to enter into an editing mode but also disagrees with the Applicant that the technique provided here cannot be performed as a mental task. A human may listen for a verbal command indicating that an edit is to be made and then decide to make the edit when it is further determined that the user who issued the verbal command is looking directly at a particular word/phrase/sentence/region. This determination that the user is looking directly at a word/phrase/sentence/region is an indication that the speed at which the user’s eyes are darting across the page (where the word/phrase/sentence/region) has been greatly reduced, since the user isn’t reading through the entire text but is instead focused on a region. A human may apply these two conditions to determine that an edit should be made. The Applicant continues on (Remarks: page 13 par 3 – page 14) to indicate that the subject matter of claim 1 recites additional elements demonstrating that the claim as a whole integrates the exception into a practical application, the practical application being an improvement in the efficiency of dictating and editing text with an electronic device and a digital assistant. The Examiner holds that this determination of whether a user intends to enter a dictation mode or editing mode is itself a mental process, and the steps taken to make the determination are also mental processes. The limitations which the Applicant refers to in order to indicate that additional elements are recited, are by themselves, mental processes. The processing of a received verbal command, and the checking of if a user’s gaze has a speed lower than a predetermined speed level, as indicated earlier, are mental processes which a human may perform through listening, observing and mentally making a decision based on the listening and observing. These tasks are by themselves mental tasks and cannot be relied upon as additional elements. The Applicant then indicates (Remarks: page 15 par 2) that ‘the claims amount to significantly more than any alleged abstract idea as they recite a specific machine learning model trained in a specific manner to perform a specific task.’ The provided machine learning model here is indicated as being trained to ‘determine the mode based on the detected gaze of the user …’ In this situation, the claims recite a machine learning model that is trained to perform, as described earlier, mental processes. The limitation does not indicate how the machine learning was actually trained, or particular machine learning models that are being used. As presented, the claims recite a generic machine learning being applied to perform mental processes and this is equivalent to a human simply applying reasoning and decision-making, which can then further be interpreted in this sense as simply having a generic computer perform a desired task. The Examiner hereby maintains the 35 U.S.C. 101 rejection. Regarding the 35 U.S.C. 103 rejection given to the claims, the Applicant has amended the independent claims to move them away from the previous presentation. The Applicant indicates (Remarks: page 16 par 2) that the applied prior art is silent on a machine learning model that considers both the velocity of a gaze on a word and an utterance, to determine to enter an editing mode. The applicant also argues against the use of the Thörn reference which considers a dwell time of the user’s gaze on an activation area and ‘does not consider “velocity of detected gaze of the user” at all.’ The Examiner notes that, the velocity of the gaze can be computed from the dwell time, the dwell time being higher when the gaze velocity reduces, showing that the gaze velocity can be inferred from a computation of the dwell time, leading the Examiner to continue to apply the reference of Thörn to address this limitation. Applicant’s arguments with respect to the independent claims have been considered but are moot due to the new ground of rejection necessitated by the amendment to the claims. The claims will be considered by their current presentation in its appropriate section below. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1, 3 – 9, 11 – 37 are rejected under 35 U.S.C. 101 because this claimed invention is directed to a judicial exception without significantly more. Independent claims 1, 19 and 20 provide teaching for detecting a gaze of a user, applying a machine learning model to the gaze to determine if to enter a dictation mode, when in a dictation mode, receiving an utterance, and applying a machine learning model determine that a user uttered a utterance that includes a predetermined word and a gaze of the user over a displayed word has a velocity below a predetermined value to determine to enter an editing mode, after determining to enter an editing mode, determining a change to be made to the displayed word and applying the change to edit the word, and then when it is determined that an editing mode shouldn’t be entered, displaying a text associated with the utterance on a screen of the electronic device. Nothing in the claims preclude the claimed technique from being performed in the human mind. The entire process involves data gathering through collecting a user’s gaze information and the collection of an utterance; data analysis through determining if to enter a dictation mode, a determination that the user’s gaze has velocity below a threshold, a determination to enter an editing mode, a determination to make an edit, and also a determination to display a textual representation based on a decision not to enter an editing mode; data transformation by converting the utterance into text, changing one word into another word through an edit process; and data presentation through the displaying of a textual representation of the utterance on a screen. A human may observe a user’s gaze and apply this as an indication that the user intends on beginning a dictation, collect utterances from the user who has signalled through a gaze that a dictation should begin, receive the user’s utterance, observe that the user slows a gaze on a particular word and that the user also utters a command that is an indication that the user wants to make an edit, in order to determine if to begin to make edits, when it is determined that an edit needs to be made, decide on the word change that needs to be made and actually making the change, and when it is determined that there is no need to make edits, display a transcription of the utterance for the user to observe. The claim hereby recites a mental process. This judicial exception is not integrated into a practical application as the claims simply teach of data gathering, analysis, transformation, and presentation. While the claims make mention of a storage medium, processors, an electronic device, these are recited in generic terms. The invention is not tied to any particular defining structure and simply provides instructions to apply the judicial exception. The technique can be performed by a generic computer which would be presented as a tool to implement the abstract idea (classifiable as automation of the mental process steps). The Specification in [0048] provides several computer devices suitable to read upon the limitations of this claim. The computer parts are recited at a high level of generality that they amount to no more than mere instructions to apply the exception using a generic computer. The trained machine learning model serves as an additional element used to make certain decisions, the decisions being to either enter a dictation mode or an editing mode. A generic machine learning model is provided in [0200] of the Specification, and it is recited without specificity on which machine learning model is applied, nor the specificities on how to apply the machine learning model. The presence of a machine learning model serves the purpose of providing nothing more than mere instructions to implement an abstract idea on a generic computer. A machine learning model is used to generally apply the abstract idea without limiting how the trained machine learning model functions. The machine learning model is described at high level of generality that mentioning it amounts to using a computer with a generic machine learning model to apply the abstract idea. A human who understands how to determine to activate a dictation mode and an editing mode based on predetermined gazes could perform the steps of this claim. The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception because the invention is not tied to a practical application. The claims provide techniques that amount to no more than mere instructions that apply the judicial exception which can be performed by a generic device. While the claims make mention of a trained machine learning model, the claims do not recite specifics on how the model is performed, and therefore still does not amount to significantly more than the mentioned judicial exception. Mere instructions to apply an exception using a generic device cannot provide an inventive concept. Claims 1, 19 and 20 are not eligible. Claim 3 provides determining whether to enter the dictation mode and the editing mode based on the machine learning model. The machine learning model serves as an additional element used to make a decision. A generic machine learning model is provided in [0200] of the Specification, and it is recited without specificity on which machine learning model is applied, nor the specificities on how to apply the machine learning models. The presence of a machine learning model serves the purpose of providing nothing more than mere instructions to implement an abstract idea on a generic computer. A machine learning model is used to generally apply the abstract idea without limiting how the trained machine learning model functions. The machine learning model is described at high level of generality that mentioning it amounts to using a computer with a generic machine learning model to apply the abstract idea. A human who understands how to determine to activate a dictation mode and an editing mode could perform the steps of this claim. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 4 provides that the determination based on the gaze to enter a dictation mode involves determining whether the detected gaze is directed at a text field displayed on the screen of an electronic device, as detected by the machine learning model able to determine a mode. A human may observe the gaze of a user and if it is observed to be directed to a particular text field, encourage the user to begin dictating. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 5 provides that the determination based on the gaze to enter a dictation mode involves determining a first location on the text field where the detected gaze of the user is directed to, as detected by the machine learning model able to determine a mode. A human may observe the gaze of a user and if it is observed to be directed to a particular location of a text field, encourage the user to begin dictating. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 6 provides teaching for determining whether the detected gaze is directed at a text field displayed on a screen comprises determining a time the gaze was directed, and determining that the gaze time exceeds a threshold. A human may observe a user’s gaze and time the gaze to collect the duration of the gaze. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 7 provides applying the gaze and the utterance to determine whether to enter the editing mode comprises determining that the user’s gaze is directed to a second location on the text field and a determination is made not to enter the editing mode when a determination that the second location is at the end of the text displayed in the text field, as detected by the machine learning model able to determine a mode. A human may receive a user’s utterance and observe the user’s gaze to be at a particular location at the end of a text field, and determine based on the observation that an edit does not need to be made. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 8 provides applying the gaze and the utterance to determine to enter an editing mode based on the location being a word of text displayed in the text field, as detected by the machine learning model able to determine a mode. If a human observes that the user is gazing at a word of text in the text field, the human may determine that as a reason to activate editing. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 9 provides applying the gaze and the utterance to determine to enter the editing mode based on a determination that a gaze spread of the detected gaze is below a spread threshold, as detected by the machine learning model able to determine a mode. A human could observe the range of a user’s eye focus on certain displayed words, and if the range isn’t wide, up to a threshold, the human may activate editing. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 11 provides determining a third location on a screen to display the textual representation. A human may select a different new location to write out a text transcription. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 12 provides that the third location is based on the user’s gaze on the screen. A human may observe a user’s gaze to see the location the user would like to place the transcription, and place the transcription at the location the user looked at. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 13 provides that the third location is based on the end of the text displayed on the screen. A human may continue to place a transcription at the end of the previous transcription text, so as to continue writing. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 14 provides that associated with a determination to enter an editing mode, a word displayed on a screen is determined to be edited. A human may, after activating editing, determine the presented word that is to be edited. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 15 provides that the determination of the displayed word to edit is based on one or more of a detected gaze, distance between a location of the detected gaze of the user and the word, a dwell time of the detected gaze and the utterance. A human may determine a word to edit based on observation of the user to determine one of the gaze direction, distance between the user’s gaze and the word, dwell time of the detected gaze, or based on listening to the user’s utterance. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 16 provides that the determined change to the word is based on the utterance and a context of displayed words. A human may listen to the user’s utterance and determine the context of the displayed words to determine the change to be made to the word. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 17 provides teaching for the determining the word to edit based on a linguistic property of the word. A human may read through the words of a sentence, consider the linguistic property of the word, and determine to edit a word that is out of place based on its linguistic property. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 18 provides the determination of a first distance between a location of a detected gaze and a first word to be equal to a second distance between a location of the detected gaze and a second word, and if these are equal, a determination is made to edit the first and second words. This is a calculation that a human may perform by observing the user’s gaze location to be equally between two words, or even with the user indicating the location of his/her gaze, so that the human may activate the editing of both words. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 21 provides that the machine learning model is trained to determine a mode based on different factors which include a direction of a user’s gaze, a proximity of the detected gaze to a word displayed on the screen, and a dwell time of the gaze. A human may make a decision based on predetermined factors associated with a gaze such as determining if the user performs any of the above-stated operations. The indicated machine learning model is simply provided here in generic terms as a tool applied to perform the intended determination action. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 22 provides determining whether to enter the dictation mode and the editing mode based on the machine learning model. The machine learning model serves as an additional element used to make a decision. A generic machine learning model is provided in [0200] of the Specification, and it is recited without specificity on which machine learning model is applied, nor the specificities on how to apply the machine learning models. The presence of a machine learning model serves the purpose of providing nothing more than mere instructions to implement an abstract idea on a generic computer. A machine learning model is used to generally apply the abstract idea without limiting how the trained machine learning model functions. The machine learning model is described at high level of generality that mentioning it amounts to using a computer with a generic machine learning model to apply the abstract idea. A human who understands how to determine to activate a dictation mode and an editing mode could perform the steps of this claim. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 23 provides that the determination based on the gaze to enter a dictation mode involves determining whether the detected gaze is directed at a text field displayed on the screen of an electronic device, as detected by the machine learning model able to determine a mode. A human may observe the gaze of a user and if it is observed to be directed to a particular text field, encourage the user to begin dictating. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 24 provides applying the gaze and the utterance to determine whether to enter the editing mode comprises determining that the user’s gaze is directed to a second location on the text field and a determination is made not to enter the editing mode when a determination that the second location is at the end of the text displayed in the text field, as detected by the machine learning model able to determine a mode. A human may receive a user’s utterance and observe the user’s gaze to be at a particular location at the end of a text field, and determine based on the observation that an edit does not need to be made. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 25 provides applying the gaze and the utterance to determine to enter the editing mode based on a determination that a gaze spread of the detected gaze is below a spread threshold, as detected by the machine learning model able to determine a mode. A human could observe the range of a user’s eye focus on certain displayed words, and if the range isn’t wide, up to a threshold, the human may activate editing. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 26 provides determining a third location on a screen to display the textual representation. A human may select a different new location to write out a text transcription. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 27 provides that associated with a determination to enter an editing mode, a word displayed on a screen is determined to be edited. A human may, after activating editing, determine the presented word that is to be edited. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 28 provides the determination of a first distance between a location of a detected gaze and a first word to be equal to a second distance between a location of the detected gaze and a second word, and if these are equal, a determination is made to edit the first and second words. This is a calculation that a human may perform by observing the user’s gaze location to be equally between two words, or even with the user indicating the location of his/her gaze, so that the human may activate the editing of both words. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 29 provides that the machine learning model is trained to determine a mode based on different factors which include a direction of a user’s gaze, a proximity of the detected gaze to a word displayed on the screen, and a dwell time of the gaze. A human may make a decision based on predetermined factors associated with a gaze such as determining if the user performs any of the above-stated operations. The indicated machine learning model is simply provided here in generic terms as a tool applied to perform the intended determination action. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 30 provides determining whether to enter the dictation mode and the editing mode based on the machine learning model. The machine learning model serves as an additional element used to make a decision. A generic machine learning model is provided in [0200] of the Specification, and it is recited without specificity on which machine learning model is applied, nor the specificities on how to apply the machine learning models. The presence of a machine learning model serves the purpose of providing nothing more than mere instructions to implement an abstract idea on a generic computer. A machine learning model is used to generally apply the abstract idea without limiting how the trained machine learning model functions. The machine learning model is described at high level of generality that mentioning it amounts to using a computer with a generic machine learning model to apply the abstract idea. A human who understands how to determine to activate a dictation mode and an editing mode could perform the steps of this claim. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 31 provides that the determination based on the gaze to enter a dictation mode involves determining whether the detected gaze is directed at a text field displayed on the screen of an electronic device, as detected by the machine learning model able to determine a mode. A human may observe the gaze of a user and if it is observed to be directed to a particular text field, encourage the user to begin dictating. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 32 provides applying the gaze and the utterance to determine whether to enter the editing mode comprises determining that the user’s gaze is directed to a second location on the text field and a determination is made not to enter the editing mode when a determination that the second location is at the end of the text displayed in the text field, as detected by the machine learning model able to determine a mode. A human may receive a user’s utterance and observe the user’s gaze to be at a particular location at the end of a text field, and determine based on the observation that an edit does not need to be made. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 33 provides applying the gaze and the utterance to determine to enter the editing mode based on a determination that a gaze spread of the detected gaze is below a spread threshold, as detected by the machine learning model able to determine a mode. A human could observe the range of a user’s eye focus on certain displayed words, and if the range isn’t wide, up to a threshold, the human may activate editing. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 34 provides determining a third location on a screen to display the textual representation. A human may select a different new location to write out a text transcription. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 35 provides that associated with a determination to enter an editing mode, a word displayed on a screen is determined to be edited. A human may, after activating editing, determine the presented word that is to be edited. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 36 provides the determination of a first distance between a location of a detected gaze and a first word to be equal to a second distance between a location of the detected gaze and a second word, and if these are equal, a determination is made to edit the first and second words. This is a calculation that a human may perform by observing the user’s gaze location to be equally between two words, or even with the user indicating the location of his/her gaze, so that the human may activate the editing of both words. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim 37 provides that the machine learning model is trained to determine a mode based on different factors which include a direction of a user’s gaze, a proximity of the detected gaze to a word displayed on the screen, and a dwell time of the gaze. A human may make a decision based on predetermined factors associated with a gaze such as determining if the user performs any of the above-stated operations. The indicated machine learning model is simply provided here in generic terms as a tool applied to perform the intended determination action. This does not integrate any practical application nor does it provide any additional element sufficient to amount to more than the mentioned judicial exception. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 3, 4, 5, 6, 7, 8, 11, 12, 13, 14, 15, 19, 20, 21, 22 23 24 26 27 29 30, 31, 32, 34, 35 and 37 are rejected under 35 U.S.C. 103 as being unpatentable over Pu et al. (US 2022/0284904 A1: hereafter — Pu) in view of Sengupta, Korok, et al. (“Leveraging error correction in voice-based text entry by Talk-and-Gaze.” Proceedings of the 2020 CHI conference on human factors in computing systems. 2020: hereafter — Septunga) and further in view of Thörn (US 2015/0364140 A1). For claim 1, Pu discloses a non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device (Pu: [0239] — non-transitory storage medium; [0232] — a processor; [0233] — execution of computer instructions; [0085] — electronic device), the one or more programs including instructions for: detecting a gaze of a user (Pu: [0008] — ‘the assistant system may use gaze as an additional signal to determine when the user wants to input text and/or make an edit to the inputted text’); determining, using a machine learning model trained to determine a mode based on the detected gaze of the user, whether to enter a dictation mode (Pu: [0008] — ‘the assistant system may use gaze as an additional signal to determine when the user wants to input text and/or make an edit to the inputted text’; [0133] — machine learning being used for face tracking as well as for recognising gestures (indicating the use of machine learning for determining to enter a dictation mode using a detected gaze)); in accordance with a determination to enter the dictation mode: receiving an utterance (Pu: [0187] — ‘[a]s an example and not by way of limitation, when the user focuses on a field with gaze, the assistant system 140 may prompt the user to dictate their utterance to enter text into that field’ (the gaze being detected to prompt the user to dictate an utterance)); determining, based on the utterance in conjunction with the detected gaze of the user using the machine learning model trained to determine the mode based on the detected gaze of the user, whether to enter an editing mode, wherein determining to enter the editing mode comprises determining that the utterance includes one or more predetermined words and that a velocity of the detected gaze of the user over a word displayed on a screen of the electronic device is below a predetermined threshold (Pu: [0187] — ‘[i]f the user indicates they want to make an edit, the user's gaze on a particular section of text may be used as a signal by the assistant system 140 to determine what to prompt the user to edit’ (user gaze to determine to enter an editing mode); [0133] — machine learning being used for face tracking as well as for recognising gestures (indicating the use of machine learning for determining to enter an editing mode using a detected gaze); [0213] — the user making use of gaze to make edits to a word by focusing a gaze on a word); in accordance with a determination to enter the editing mode: determining a change to be made to the word displayed on the screen of the electronic device (Pu: [0210], FIGs. 9C, 9D — ‘In FIG. 9C, the circle 915 may indicate the user's gaze input, which is fixated at the block “in twenty” 910. This means the user wants to edit the block “in twenty” 910. FIG. 9D illustrates an example user interface showing an edit to the block.’ (choosing to edit ‘in twenty’)); [0210], FIG. 9D — choosing the replacement as ‘in thirty’); and editing the word by applying the change to the word (Pu: [0210], FIG. 9E — enacting the change to ‘in thirty’); and in accordance with a determination not to enter the editing mode, displaying a textual representation of the utterance on the screen of the electronic device. The reference of Pu provides teaching for a applying a user’s gaze to determine to enter an editing mode. This reference however differs from the claimed invention in that the fails to teach of determining to enter an editing mode based on a combination of both a predetermined word utterance and a user gaze. This teaching however isn’t new to the art as the reference of Sengupta is now introduced to teach this as: determining, based on the utterance in conjunction with the detected gaze of the user using the machine learning model trained to determine the mode based on the detected gaze of the user, whether to enter an editing mode, wherein determining to enter the editing mode comprises determining that the utterance includes one or more predetermined words and that a velocity of the detected gaze of the user over a word displayed on a screen of the electronic device is below a predetermined threshold (Sengupta: page 2 Col 1 par 2 — a technique called V-TaG which selects an erroneous word when a user utters a voice command while gazing at the word/location; page 4 Col 2 2. — ‘focusing on the incorrect word and then saying “select” to select the erroneous word’ (teaching of providing a predetermined word for a user to utter)). Hence, before the effective filing date of the claimed invention, one of ordinary skill in the art would have found it obvious to improve upon the use of just a user’s gaze to enter an editing mode as taught by the reference Pu, by applying the known technique of Sengupta which applies both a user’s gaze and a voice command, to determine to enter an editing mode, to thereby come up with the claimed invention. The combination of both prior art elements would have provided the predictable result of reducing accidental edits that might be determined based on gaze alone, whereby applying both gaze and a speech command would provide a proper confirmation of the user’s intent to make an edit. See KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398, 415-421, 82 USPQ2d 1385, 1395-97 (2007). The combination of Pu in view of Sengupta provides teaching for determining to enter an edit mode based on both a user’s detected gaze and on a speech command, but differs from the claimed invention in that the claimed invention further provides that the gaze determination requires the user’s gaze to be below a predetermined velocity threshold. This is however not new to the art as the reference of Thorn is now introduced to teach this as: determining, based on the utterance in conjunction with the detected gaze of the user using the machine learning model trained to determine the mode based on the detected gaze of the user, whether to enter an editing mode, wherein determining to enter the editing mode comprises determining that the utterance includes one or more predetermined words and that a velocity of the detected gaze of the user over a word displayed on a screen of the electronic device is below a predetermined threshold (Thorn: [0069] — activating text editing of speech-to-text generated text, through user gaze, based on detecting a dwell time of the user’s gaze on the activation area, to know if a dwell time exceeds a threshold (thereby teaching of determining that a user focuses the gaze at a particular region with a slow speed/velocity, obtained based on the dwell time, the dwell time being higher when the gaze velocity reduces, showing that the gaze velocity can be obtained from the computation of the dwell time)). Hence, before the effective filing date of the claimed invention, one of ordinary skill in the art would have found it obvious to improve upon the activation of an editing mode based on gaze detection on a word meant to be edited as determined through machine learning just as taught by the combination of Pu in view of Sengupta, by incorporating the detection of a dwell time of the gaze for entering the editing mode as taught by the reference of Thorn, to thereby come up with the claimed invention. The combination of both prior art elements would have provided the predictable result of applying the determined speed of the user’s gaze on a word intended to be edited as a confirmation that the user does indeed desire to edit that word over other words which the user’s gaze quickly dashes over, leading to a simple way of detecting the desire to make an edit. See KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398, 415-421, 82 USPQ2d 1385, 1395-97 (2007). For claim 3, the combination of Pu in view of Sengupta further in view of Thorn discloses the non-transitory computer-readable storage medium, wherein determining whether to enter the dictation mode and determining whether to enter the editing mode are determined by the same machine learning model (Pu: [0133] — machine learning being used for face tracking as well as for recognising gestures (indicating the use of machine learning for determining to enter a dictation mode and to enter an editing mode using a detected gaze)). For claim 4, claim 1 is incorporated and the combination of Pu in view of Sengupta further in view of Thorn discloses the non-transitory computer-readable storage medium, wherein determining, using a machine learning model trained to determine a mode based on the detected gaze of the user, whether to enter the dictation mode further comprises: determining whether the detected gaze of the user is directed at a text field displayed on a screen of the electronic device (Pu: [0187] — determining that a user should begin a dictation in a text field when a user gaze is detected at that text field of a display on a user interface). For claim 5, claim 4 is incorporated and the combination of Pu in view of Sengupta further in view of Thorn discloses the non-transitory computer-readable storage medium wherein determining, using a machine learning model trained to determine a mode based on the detected gaze of the user, whether to enter a dictation mode further comprises: determining a first location on the text field where the detected gaze of the user is directed (Pu: [0187] — having the user focus a gaze attention on a source, or on an assistant icon (to indicate the fixing of the gaze attention on a location of the screen)). For claim 6, claim 4 is incorporated and the combination of Pu in view of Sengupta further in view of Thorn discloses the non-transitory computer-readable storage medium, wherein determining whether the detected gaze of the user is directed at the text field displayed on a screen of the electronic device (Pu: [0187] — determining that a user should begin a dictation in a text field when a user gaze is detected at that text field of a display on a user interface) further comprises: determining a time that the detected gaze of the user is directed at the text field (Thorn: [0069] — activating text editing of speech-to-text generated text, through user gaze, based on detecting a dwell time of the user’s gaze on the activation area, to know if a dwell time exceeds a threshold); and determining that the detected gaze of the user is directed at the text field in accordance with a determination that the time exceeds a threshold (Thorn: [0069] — activating text editing of speech-to-text generated text, through user gaze, based on detecting a dwell time of the user’s gaze on the activation area, to know if a dwell time exceeds a threshold). For claim 7, claim 1 is incorporated and the combination of Pu in view of Sengupta further in view of Thorn discloses the non-transitory computer-readable storage medium, wherein determining, based on the utterance in conjunction with the detected gaze of the user using a machine learning model trained to determine a mode based on the detected gaze of the user, whether to enter the editing mode further comprises: determining a second location on the text field where the detected gaze of the user is directed (Pu: FIG. 9D Part 915; [0210] — directing the gaze at a location on the display for the purpose of editing); and determining not to enter the editing mode in accordance with a determination that the second location is at the end of text displayed in the text field (Pu: FIG. 9E Part 915; [0210] — the user’s gaze is directed to a location at the end of the text, and this is not an indication to enter an editing mode). For claim 8, claim 7 is incorporated and the combination of Pu in view of Sengupta further in view of Thorn discloses the non-transitory computer-readable storage medium, wherein determining, based on the utterance in conjunction with the detected gaze of the user using a machine learning model trained to determine a mode based on the detected gaze of the user, whether to enter the editing mode further comprises: determining to enter the editing mode in accordance with a determination that the second location is on a word of text displayed in the text field (Pu: FIG. 9D Part 915; [0210] — directing the gaze at a location on the display for the purpose of editing, the gaze being directed to the text displayed on the screen). For claim 11, claim 1 is incorporated and the combination of Pu in view of Sengupta further in view of Thorn discloses the non-transitory computer-readable storage medium, the one or more programs further including instructions for: determining a third location to display the textual representation of the utterance on the screen of the electronic device (Pu: FIG. 6D Part 615, [0207] — ‘FIG. 6D illustrates an example user interface showing the new dictation. The user may say “I'll be there in thirty 615”’ (showing a new location 615 to display the textual representation)). For claim 12, claim 11 is incorporated and the combination of Pu in view of Sengupta further in view of Thorn discloses the non-transitory computer-readable storage medium, wherein the third location to display the textual representation of the utterance on the screen of the electronic device is determined based on the location of the user’s gaze on the screen (Pu: [0187] — determining that a user should begin a dictation in a text field when a user gaze is detected at that text field of a display on a user interface (the text being placed in the gaze location which is the text field region)). For claim 13, claim 11 is incorporated and the combination of Pu in view of Sengupta further in view of Thorn discloses the non-transitory computer-readable storage medium, wherein the third location to display the textual representation of the utterance on the screen of the electronic device is determined based on the end of text displayed on the screen of the electronic device (Pu: FIG. 6D Part 615 — the location of the displayed textual representation of the utterance comes shows up at the end of the previous text, after the text in Part 500). For claim 14, claim 1 is incorporated and the combination of Pu in view of Sengupta further in view of Thorn discloses the non-transitory computer-readable storage medium, the one or more programs further including instructions for: in accordance with a determination to enter the editing mode, determining the word displayed on the screen of the electronic device to edit (Pu: [0210], FIGs. 9C, 9D — ‘In FIG. 9C, the circle 915 may indicate the user's gaze input, which is fixated at the block “in twenty” 910. This means the user wants to edit the block “in twenty” 910. FIG. 9D illustrates an example user interface showing an edit to the block.’ (choosing to edit ‘in twenty’)). For claim 15, claim 14 is incorporated and the combination of Pu in view of Sengupta further in view of Thorn discloses the non-transitory computer-readable storage medium, wherein determining the word displayed on the screen of the electronic device to edit is based on one or more of the detected gaze of the user, the distance between a location of the detected gaze of the user and the word, a dwell time of the detected gaze of the user, and the utterance (Pu: [0210], FIG. 9C Part 915 — “In FIG. 9C, the circle 915 may indicate the user’s gaze input, which is fixated at the block “in twenty” 910” as an indication of determining the word to be edited). As for claim 19, electronic device of claim 19 and computer programme product claim 1 are related as device system for performing the available programme instructions. Pu in [0232] provides a processor and in [0239] provides storage memory, suitable to read upon the limitations of this claim. Accordingly, claim 19 is similarly rejected under the same rationale as applied above with respect to computer programme product claim 1. As for claim 20, method claim 20 and computer-readable medium claim 1 are related as method detailing procedures taken to implement the available computer-programmable instructions. Pu in FIG. 16 provides a method, an electronic device in [0085], a processor in [0232] and a storage medium in [0239], suitable to address the limitations of this claim. Accordingly, claim 20 is similarly rejected under the same rationale as applied above with respect to computer-readable medium claim 1. For claim 21, claim 1 is incorporated and the combination of Pu in view of Sengupta further in view of Thorn discloses the non-transitory computer-readable storage medium wherein the machine learning model is trained to determine a mode based on a plurality of factors, the plurality of factors including at least one of a direction of the detected gaze of the user (Pu: [0203] — determines that a user’s eyes are moving in a direction), a proximity of the detected gaze of the user to a word displayed on the screen of the electronic device, and a dwell time of the detected gaze of the user (Thorn: [0069] — determining user gaze dwell time). As for claim 22, electronic device of claim 22 and computer programme product claim 3 are related as device system for performing the available programme instructions. Accordingly, claim 22 is similarly rejected under the same rationale as applied above with respect to computer programme product claim 3. As for claim 23, electronic device of claim 23 and computer programme product claim 4 are related as device system for performing the available programme instructions. Accordingly, claim 23 is similarly rejected under the same rationale as applied above with respect to computer programme product claim 4. As for claim 24, electronic device of claim 24 and computer programme product claim 7 are related as device system for performing the available programme instructions. Accordingly, claim 24 is similarly rejected under the same rationale as applied above with respect to computer programme product claim 7. As for claim 26, electronic device of claim 26 and computer programme product claim 11 are related as device system for performing the available programme instructions. Accordingly, claim 26 is similarly rejected under the same rationale as applied above with respect to computer programme product claim 11. As for claim 27, electronic device of claim 27 and computer programme product claim 14 are related as device system for performing the available programme instructions. Accordingly, claim 27 is similarly rejected under the same rationale as applied above with respect to computer programme product claim 14. As for claim 29, electronic device of claim 29 and computer programme product claim 21 are related as device system for performing the available programme instructions. Accordingly, claim 29 is similarly rejected under the same rationale as applied above with respect to computer programme product claim 21. As for claim 30, method claim 30 and computer-readable medium claim 3 are related as method detailing procedures taken to implement the available computer-programmable instructions. Accordingly, claim 30 is similarly rejected under the same rationale as applied above with respect to computer-readable medium claim 3. As for claim 31, method claim 31 and computer-readable medium claim 4 are related as method detailing procedures taken to implement the available computer-programmable instructions. Accordingly, claim 31 is similarly rejected under the same rationale as applied above with respect to computer-readable medium claim 4. As for claim 32, method claim 32 and computer-readable medium claim 7 are related as method detailing procedures taken to implement the available computer-programmable instructions. Accordingly, claim 32 is similarly rejected under the same rationale as applied above with respect to computer-readable medium claim 7. As for claim 34, method claim 34 and computer-readable medium claim 1 are related as method detailing procedures taken to implement the available computer-programmable instructions. Accordingly, claim 34 is similarly rejected under the same rationale as applied above with respect to computer-readable medium claim 1. As for claim 35, method claim 35 and computer-readable medium claim 14 are related as method detailing procedures taken to implement the available computer-programmable instructions. Accordingly, claim 35 is similarly rejected under the same rationale as applied above with respect to computer-readable medium claim 14. As for claim 37, method claim 37 and computer-readable medium claim 21 are related as method detailing procedures taken to implement the available computer-programmable instructions. Accordingly, claim 37 is similarly rejected under the same rationale as applied above with respect to computer-readable medium claim 21. Claim 9, 25 and 33 are rejected under 35 U.S.C. 103 as being unpatentable over Pu (US 2022/0284904 A1) in view of Sengupta (“Leveraging error correction in voice-based text entry by Talk-and-Gaze.” Proceedings of the 2020 CHI conference on human factors in computing systems. 2020) further in view of Thorn (US 2015/0364140 A1) as applied to claims 1, 19 and 20, further in view of Schwartz (US 2020/0168038 A1). For claim 9, claim 1 is incorporated and the combination of Pu in view of Sengupta further in view of Thorn discloses the non-transitory computer-readable storage medium, wherein determining, based on the utterance in conjunction with the detected gaze of the user using a machine learning model trained to determine a mode based on the detected gaze of the user, whether to enter the editing mode further comprises: determining to enter the editing mode [[in accordance with a determination that a gaze spread of the detected gaze is below a spread threshold]] (Pu: [0210] — ‘FIG. 9C illustrates an example user interface showing a gaze input. In FIG. 9C, the circle 915 may indicate the user's gaze input, which is fixated at the block "in twenty" 910. This means the user wants to edit the block "in twenty" 910.’ (teaching of a user using a gaze to have the system fixate on a certain area, the fixation on the words leading to a determination to perform editing)). The combination of Pu in view of Sengupta further in view of Thorn provides teaching for activating an edit by determining where a user’s gaze is fixated upon, but differs from the claimed invention in that the claimed invention further provides teaching for a determination that a gaze spread is below a threshold. This isn’t new to the art as the reference of Schwartz is now introduced to teach this as: determining to enter the editing mode in accordance with a determination that a gaze spread of the detected gaze is below a spread threshold (Schwartz: Claim 7 — determining a high engagement level when a gaze direction is determined to be within a pre-determined are of focus (a gaze direction being within an area of focus indicates that the gaze spread is within a threshold, and the high engagement level indicates a fixation on the gazed-upon area)). Hence, before the effective filing date of the claimed invention, one of ordinary skill in the art would have found it obvious to improve upon the teaching of the combination of Pu in view of Sengupta further in view of Thorn which teach of determining where a gaze is fixated upon to determine that an edit should be made, by applying the known technique of Schwartz which determines that a user’s gaze is fixed on a location through determining that there is a high engagement level, through determining that the area of focus of the gaze is less than a threshold, to thereby come up with the claimed invention. The combination of both prior art elements would have provided the predictable result of being able to properly determine the user’s focus location without confusing it with a larger area, so that the editing can be localised to only the words that need editing. See KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398, 415-421, 82 USPQ2d 1385, 1395-97 (2007). As for claim 25, electronic device of claim 25 and computer programme product claim 9 are related as device system for performing the available programme instructions. Accordingly, claim 25 is similarly rejected under the same rationale as applied above with respect to computer programme product claim 9. As for claim 33, method claim 33 and computer-readable medium claim 9 are related as method detailing procedures taken to implement the available computer-programmable instructions. Accordingly, claim 33 is similarly rejected under the same rationale as applied above with respect to computer-readable medium claim 9. Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Pu (US 2022/0284904 A1) in view of Sengupta (“Leveraging error correction in voice-based text entry by Talk-and-Gaze.” Proceedings of the 2020 CHI conference on human factors in computing systems. 2020) further in view of Thorn (US 2015/0364140 A1) as applied to claim 14, further in view of Goel at al. (US 2015/0242391 A1: hereafter — Goel). For claim 16, claim 1 is incorporated and the combination of Pu in view of Sengupta further in view of Thorn discloses the non-transitory computer-readable storage medium, wherein determining the change to be made to the word displayed on the screen of the electronic device is based on the utterance [[and a context of the words displayed on the screen of the electronic device]] (Pu: [0210] — ‘As illustrated in FIG. 9D, the user may have dictated the edit as “in thirty” 920 to replace “in twenty” 910’ (user dictates the utterance that is to be used to make the edit)). The combination of Pu in view of Sengupta further in view of Thorn however fails to teach the further limitation of this claim regarding the use of context, for which the reference of Goel is now introduced to teach as: the non-transitory computer-readable storage medium, wherein determining the change to be made to the word displayed on the screen of the electronic device is based on the utterance and a context of the words displayed on the screen of the electronic device (Goel: [0041] — providing replacements for one or more words by making use of context-appropriate words; FIG. 3 — making corrections to words displayed on a screen of an electronic device). Hence, before the effective filing date of the claimed invention, one of ordinary skill in the art would have found it obvious to combine the known teaching of Goel which determines the change to be made to the word based on a context of displayed words, with the teaching of determining the change to be made based on the utterance as taught by the combination of Pu in view of Sengupta further in view of Thorn, to thereby come up with the claimed invention. The combination of both prior art elements would have provided the predictable result of ensuring that the replacement word is contextually appropriate for the sentence it is being inserted into, leading to a grammatically correct sentence. See KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398, 415-421, 82 USPQ2d 1385, 1395-97 (2007). Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Pu (US 2022/0284904 A1) in view of Sengupta (“Leveraging error correction in voice-based text entry by Talk-and-Gaze.” Proceedings of the 2020 CHI conference on human factors in computing systems. 2020) further in view of Thorn (US 2015/0364140 A1) as applied to claim 14, further in view of Bojja et al. (US 2017/0185581 A1: hereafter — Bojja). For claim 17, claim 14 is incorporated but the combination of Pu in view of Sengupta further in view of Thorn fails to disclose the teaching of this claim, for which the reference of Bojja is now introduced to teach as the non-transitory computer-readable storage medium, wherein determining the word displayed on the screen of the electronic device to edit is based on a linguistic property of the word (Bojja: [0036] — grammar error correction method which parses an input sentient to determine parts of speech of individual words (the part of speech being a linguistic property)). The combination of Pu in view of Sengupta further in view of Thorn provides teaching for determining a change is to be made to a word displayed on a screen, but differs from the claimed invention in that the claimed invention further provides teaching for determining the word based on a linguistic property of the word. This isn’t new to the art as the reference of Bojja is seen to teach above, the linguistic property being a part-of-speech. Hence, before the effective filing date of the claimed invention, one of ordinary skill in the art would have found it obvious to improve upon the teaching of the combination of Pu in view of Sengupta further in view of Thorn which teaches of determining that a change is to be made to a word displayed on a screen, by applying the known technique of Bojja which determines the part-of-speech of each word in a sentence for the purpose of performing error correction, to thereby come up with the claimed invention. The combination of both prior art elements would have provided the predictable result of ensuring that the arrangement of the words in the sentence are linguistically correct, so the discovery of a linguistic misalignment would signal an error in the sentence. See KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398, 415-421, 82 USPQ2d 1385, 1395-97 (2007). Claim 18, 28 and 36 are rejected under 35 U.S.C. 103 as being unpatentable over Pu (US 2022/0284904 A1) in view of Sengupta (“Leveraging error correction in voice-based text entry by Talk-and-Gaze.” Proceedings of the 2020 CHI conference on human factors in computing systems. 2020) further in view of Thorn (US 2015/0364140 A1) as applied to claims 1, 19 and 20, further in view of POWDERLY et al. (US 2018/0307303 A1: hereafter — Powderly). For claim 18, claim 1 is incorporated but the combination of Pu in view of Sengupta further in view of Thorn fails to disclose the limitations of this claim, for which the reference of Powderly is now introduced to teach as the non-transitory computer-readable storage medium, the one or more programs further including instructions for: in accordance with a determination that a first distance between a location of the detected gaze of the user and a first word is equal to a second distance between a location of the detected gaze of the user and a second word, determining to edit both the first word and the second word (Powderly: [0308] — a user may select a phrase using gaze by focusing on first word and a second word, so that both words (and those in-between) may be selected for editing (the centre of the user’s gaze would be a point equidistant to both the first and second words)). The combination of Pu in view of Sengupta further in view of Thorn provides teaching for obtaining a transcription of words of an utterance of a user received based on a user’s gaze, but differs from the claimed invention in that the claimed invention further provides teaching for determining to edit two words for which the distance between the user’s gaze and each of the words is equal. This isn’t new to the art as is seen to be taught by the reference of Powderly above. Hence, before the effective filing date of the claimed invention, one of ordinary skill in the art would have found it obvious to improve upon the teaching of the combination of Pu in view of Sengupta further in view of Thorn which applies gaze to begin a dictation and provide a transcription of the dictation from a user, by applying the known technique of Powderly which is able to make a selection of two words for an edit based on the user’s gaze having gone through both words, to thereby come up with the claimed invention. The combination of both prior art elements would have provided the predictable result of ensuring the editing of multiple words with a single action. See KSR Int’l Co. v. Teleflex Inc., 550 U.S. 398, 415-421, 82 USPQ2d 1385, 1395-97 (2007). As for claim 28, electronic device of claim 28 and computer programme product claim 18 are related as device system for performing the available programme instructions. Accordingly, claim 28 is similarly rejected under the same rationale as applied above with respect to computer programme product claim 18. As for claim 36, method claim 36 and computer-readable medium claim 18 are related as method detailing procedures taken to implement the available computer-programmable instructions. Accordingly, claim 36 is similarly rejected under the same rationale as applied above with respect to computer-readable medium claim 18. Conclusion Applicant’s amendment necessitated the new grounds of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. The prior art made of record and not relied upon is considered pertinent to Applicant’s disclosure. Tzvieli et al. (US 2021/0318558 A1) provides teaching for determining an eye movement velocity, determining the velocity is below a threshold, and also a machine learning model able to detect the eye movements [0169]. Any inquiry concerning this communication or earlier communications from the examiner should be directed to OLUWADAMILOLA M. OGUNBIYI whose telephone number is (571)272-4708. The examiner can normally be reached Monday – Thursday (8:00 AM – 5:30 PM Eastern Standard Time). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s Supervisor, PARAS D. SHAH can be reached at (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /OLUWADAMILOLA M OGUNBIYI/Examiner, Art Unit 2653 /Paras D Shah/Supervisory Patent Examiner, Art Unit 2653 09/11/2026
Read full office action

Prosecution Timeline

Show 9 earlier events
Feb 23, 2026
Examiner Interview Summary
Feb 27, 2026
Request for Continued Examination
Mar 02, 2026
Response after Non-Final Action
Mar 25, 2026
Non-Final Rejection mailed — §101, §103
Jun 11, 2026
Applicant Interview (Telephonic)
Jun 11, 2026
Examiner Interview Summary
Jun 22, 2026
Response Filed
Sep 15, 2026
Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12725620
AUDIO ENCODER AND DECODER USING A FREQUENCY DOMAIN PROCESSOR , A TIME DOMAIN PROCESSOR, AND A CROSS PROCESSING FOR CONTINUOUS INITIALIZATION
3y 0m to grant Granted Sep 01, 2026
Patent 12718584
SENSOR FUSION FOR COLLISION DETECTION
3y 0m to grant Granted Aug 25, 2026
Patent 12694891
METHOD, DEVICE AND COMPUTER PROGRAM FOR EMOTION RECOGNITION FROM A REAL-TIME AUDIO SIGNAL
3y 7m to grant Granted Jul 28, 2026
Patent 12640154
Stylizing Text-to-Speech (TTS) Voice Response for Assistant Systems
3y 5m to grant Granted May 26, 2026
Patent 12608427
Drill Back To Original Audio Clip In Virtual Assistant Initiated Lists And Reminders
1y 11m to grant Granted Apr 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
77%
Grant Probability
96%
With Interview (+19.4%)
2y 11m (~3m remaining)
Median Time to Grant
High
PTA Risk
Based on 315 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month