Prosecution Insights
Last updated: August 18, 2026
Application No. 18/981,176

MULTI-MODEL GESTURE TO AUDIO TRANSLATION

Non-Final OA §103
Filed
Dec 13, 2024
Examiner
AGAHI, DARIOUSH
Art Unit
2656
Tech Center
2600 — Communications
Assignee
Optum Inc.
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
151 granted / 179 resolved
+22.4% vs TC avg
Strong +30% interview lift
Without
With
+29.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
24 currently pending
Career history
206
Total Applications
across all art units

Statute-Specific Performance

§101
24.6%
-15.4% vs TC avg
§103
54.0%
+14.0% vs TC avg
§102
11.3%
-28.7% vs TC avg
§112
6.7%
-33.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 179 resolved cases

Office Action

§103
DETAILED ACTION This office action is in response to Applicant’s submission filed on 12/13/2024. Claims 1-20 are pending in the application of which Claims 1, 10, and 17 are independent and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement(s)(IDS) submitted on 3/26/2020, and 5/8/2026 have been considered by the examiner. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 4-10, 13 - 17, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Rahmani et al. (US20230085161A1)(herein " Rahmani "), and in further view of Hung-Ta Liu (US20130300650A1)(herein "Liu"). Regarding claims 1, 10 and 17 Rahmani teaches [A computer-implemented method comprising: - claim 1], [A system comprising: one or more processors; and one or more memories storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: - claim 10], and [One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising: - claim 17] (Rahmani, Par. 0085:” … machine readable instructions, which may be executed to configure processor circuitry to implement the translation server ... The machine readable instructions may be one or more executable programs or portion(s) of an executable program for execution by processor circuitry, such as the processor circuitry … software stored on one or more non-transitory computer readable storage media …”, and Par. 0089:” As mentioned above, the example operations of FIG. 5 may be implemented using executable instructions (e.g., computer and/or machine readable instructions) stored on one or more non-transitory computer and/or machine readable media such ... non-transitory machine readable medium, and … the terms “computer readable storage device” and “machine readable storage device” …”, and Par. 0105:” … may be implemented by the machine readable instructions of FIG. 7, may be stored in the mass storage device 728, in the volatile memory 714, in the non-volatile memory 716, and/or on a removable non-transitory computer readable storage medium “). receiving, by one or more processors, an image that depicts a facial expression and a hand position of a user; (Rahmani, Par. 0029:” … For example, a deaf person may generate signs in the field of view of a camera of the second user device 104. The translation app 108 of the second user device 104 sends video data representing the signs to the translation server 110. The translation server 110 applies the sign/text model to transform the video data into text. The translation server 110 applies a text-to-speech (TTS) model to translate the text into speech (audio data). Thus, the speech is spoken language corresponding to the video of sign language input at the second user device 104.”, and Par. 0051:” … The key points are based on features of the deaf person's body and movement. In some examples, the key points are based on hand shape, hand orientation, and hand location. In some examples, the key points are based on movement of one or more of the hands, face, head, torso, mouthing, and/or finger spelling.”, and Par. 0133:” … wherein the processor circuitry is to: identify a movement of a person in the video as a non-sign language movement; …”). generating, by the one or more processors and using a [[parallel feature extraction]] model of a multi-stage machine learning architecture, a set of facial features and a set of hand features from the image; (Rahmani, Par. 0051:” The inference server 308 translates the signs in the video to speech. For example, the sign/text engine circuitry 314 applies the sign/text model to transform signs in the video into text. … The key points are based on features of the deaf person's body and movement. In some examples, the key points are based on hand shape, hand orientation, and hand location. In some examples, the key points are based on movement of one or more of the hands, face, head, torso, mouthing, and/or finger spelling. In some examples, the key points are based on rhythm, cadence, eye gaze, and/or pauses. Any, all, and other features may be used to identify key points to be used in identifying a sign.”, and Par. 0052:”…the input video frame to facilitate obtaining visual features from the video. The sign/text engine circuitry 314 uses a pre-trained feature extraction model such as, for example, a machine learning model capable of analyzing live and/or streaming media (e.g., MediaPipe).”, and Par. 0133:” … wherein the processor circuitry is to: identify a movement of a person in the video as a non-sign language movement; …”) generating, by the one or more processors and using an aggregation model of the multi-stage machine learning architecture, a text prediction corresponding to the image based on the set of facial features, the set of hand features, and a set of defined terms associated with the multi-stage machine learning architecture; and (Rahmani, Par. 0052:”…the input video frame to facilitate obtaining visual features from the video. The sign/text engine circuitry 314 uses a pre-trained feature extraction model such as, for example, a machine learning model capable of analyzing live and/or streaming media (e.g., MediaPipe).”, and Par. 0055:” Implementing the sign-to-gloss translation model, the sign/text engine circuitry 314 also trains a single sign recognition model, which is a transformer-based model. In some examples, the single sign recognition model is based on ASL signs. … The single sign recognition model is trained with a number of single sign videos and key points obtained from the reference model. The trained single sign recognition model inputs single sign video key points and outputs a gloss label [defined terms] with a rank or confidence score (CS). In some examples, the confidence score is determined through a probability function applied to the output of the single sign recognition model. This probability score indicates a confidence level that the sign/text engine circuitry 314 has extracted the correct gloss label. In some examples, the CS ranges from 0 to 1, where the most confident recognition achieves CS=1.”, and Par. 0133:” … wherein the processor circuitry is to: identify a movement of a person in the video as a non-sign language movement; …”, and Par. 0133:” … wherein the processor circuitry is to: identify a movement of a person in the video as a non-sign language movement; …”). initiating, by the one or more processors, a prediction-based action based on the text prediction. (Rahmani, Par 0070:” The TTS engine circuitry 312 converts the text to audio data. In some examples, the TTS engine circuitry 312 tokenizes the text and identifies phonetic transcriptions to the words in the text. The TTS engine circuitry 312 converts the phonetic transcriptions into a sound wave (as audio data). In some examples, the TTS engine circuitry 312 uses the metadata indicative of the non-manual features of the sign language to create or adjust a volume, cadence, tone, and/or emotion of the speech. In other examples, the TTS engine circuitry 312 can use other speech synthesis operations.", and Par. 0133:” … wherein the processor circuitry is to: identify a movement of a person in the video as a non-sign language movement; …”). Rahmani does not teach, however Liu teaches parallel feature extraction (Liu, Par. 0037:”… facial expression may be a characteristic position of each of the face organs, or in combination of serial changes of the organs. The displacements among the user 1's eyebrows 40, eyes 41, ears 42, nose 43, or mouth 44 may be recognized to be the facial expressions. The facial expressions in the current examples are such as the variations of the eyebrows 40 shown in FIGS. 4A to 4C, the eyes 41 of FIGS. 5A to 5D including blinking single eye, alternately blinking eyes, or simultaneous [parallel] blinking eyes, and the variations of closing or opening mouth shown in FIGS. 6A to 6C, including combination of the expressions of opening mouth, closing mouth, and extending tongue, or the changes of shape as the user performing lip language or speaking.”, and Par. 0038:”… the illustrated facial expressions may be combined with the simultaneous motions of other facial organs. One of the facial expressions is such as closing single eye shown in FIG. 4A and FIG. 4B combined with opening mouth and closing mouth shown in FIGS. 6A through 6B.”, and Par. 0133:” … wherein the processor circuitry is to: identify a movement of a person in the video as a non-sign language movement; …”). Liu is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rahmani further in view of Liu to employ parallel feature extraction. Motivation to do so would significantly enhance accuracy and disambiguation, empowering the model to differentiate between signs that share structural similarities but diverge in meaning or context. Regarding claims 4, and 13, Rahmani, as modified above, teaches the computer-implemented method, and the system claims of 1, and 10, respectively. Rahmani, as modified above, further teaches wherein the image is one of a set of images in a video stream, the set of facial features and the set of hand features are based on a previous image associated with a previous time relative to the image, at least one of (i) the set of facial features comprises (a) a facial expression feature that identifies a sentiment classification for the user, (b) an eye movement feature that identifies a focus classification for the user, or (c) a lip position feature that identifies a lip movement classification based on a lip position change from the previous image to the image, or (ii) the set of hand features comprises a gesture recognition classification based on at least one of a hand or finger position change from the previous image to the image. (Rahmani, Par. 0051:” … applies the sign/text model to transform signs in the video into text. … For example, in some examples, the sign/text engine circuitry 314 identifies a sign in the video [set of images] by recognizing key points. The key points are based on features of the deaf person's body and movement. In some examples, the key points are based on hand shape, hand orientation, and hand location. In some examples, the key points are based on movement of one or more of the hands, face, head, torso, mouthing, and/or finger spelling. In some examples, the key points are based on rhythm, cadence, eye gaze, and/or pauses.”, and Par. 0052:” … non-manuals features, such as for example, facial expression, mouth gestures, and movement of the head, shoulders, and torso.”, and Par. 0069:” … the sign/text engine circuitry 314 may identify non-manual features in the video of the sign language (e.g., rhythm, cadence, eye gaze, facial expressions, and/or pauses) that are analogous to intonation in spoken language. The sign/text engine circuitry 314 can tag the text with metadata indicative of one or more of the non-manual features.”, and Par. 0070:” … In some examples, the TTS engine circuitry 312 uses the metadata indicative of the non-manual features of the sign language to create or adjust a volume, cadence, tone, and/or emotion [sentiment] of the speech.”, and Par. 0122:” … identify non-manual elements in the sign language; and alter at least one of a volume, cadence, tone, or emotion of the audio based on the non-manual elements.”, and Par. 0130:” … wherein the processor circuitry is to: identify non-manual elements in the performed signs; and tag the audio data with at least one of a volume, cadence, tone, or emotion based on the non-manual elements.”) Regarding claims 5, and 14, Rahmani, as modified above, teaches the computer-implemented method, and the system claims of 4, and 13, respectively. Rahmani, as modified above, further teaches wherein at least one of the sentiment classification is based on a change of the facial expression from a previous facial expression of the previous image, or the eye movement feature is based on a change in eye focus from the previous facial expression to the facial expression. (Rahmani, Par. 0122:” … identify non-manual elements in the sign language; and alter at least one of a volume, cadence, tone, or emotion of the audio based on the non-manual elements.”, and Par. 0130:” … wherein the processor circuitry is to: identify non-manual elements in the performed signs; and tag the audio data with at least one of a volume, cadence, tone, or emotion based on the non-manual elements.”) Note: Emotional states are identified from a video's non-manual features, which are primarily conveyed through subtle visual signals like Facial Expressions and eye gaze direction. Regarding claims 6, 15, and 20, Rahmani, as modified above, teaches the computer-implemented method, the system, and the non-transitory media claims of 1, 10, and 17, respectively. Rahmani, as modified above, further teaches wherein initiating the prediction-based action based on the text prediction comprises: generating, using a text-to-speech model, an audio signal based on the text prediction; and outputting, using a user interface, the audio signal. (Rahmani, Par. 0029:” … a deaf person may generate signs in the field of view of a camera of the second user device 104. The translation app 108 of the second user device 104 sends video data representing the signs to the translation server 110. The translation server 110 applies the sign/text model to transform the video data into text. The translation server 110 applies a text-to-speech (TTS) model to translate the text into speech (audio data). Thus, the speech is spoken language corresponding to the video of sign language input at the second user device 104. The translation server 110 sends the audio data to the translation app 108 on the first user device 102. The audio data is presented as speech via one or more loudspeakers of the first user device 102 to the hearing person. That is, the audio data is convertible into a sound wave.”, and Par. 0102:”One or more output devices 724 are also connected to the interface circuitry 720 of the illustrated example. The output device(s) 724 can be implemented, for example, by display devices …, and/or speaker.”) Regarding claim 7, Rahmani, as modified above, teaches the computer-implemented method claim of 6. Rahmani, as modified above, further teaches wherein: the user interface comprises one or more audio devices of a user device and the image is recorded by one or more imaging devices of the user device, or the user interface comprises one or more audio devices of a second device and the image is recorded by the one or more imaging devices of the user device. (Rahmani, Par. 0029:” … a deaf person may generate signs in the field of view of a camera of the second user device 104. The translation app 108 of the second user device 104 sends video data representing the signs to the translation server 110. The translation server 110 applies the sign/text model to transform the video data into text. The translation server 110 applies a text-to-speech (TTS) model to translate the text into speech (audio data). Thus, the speech is spoken language corresponding to the video of sign language input at the second user device 104. The translation server 110 sends the audio data to the translation app 108 on the first user device 102. The audio data is presented as speech via one or more loudspeakers of the first user device 102 to the hearing person. That is, the audio data is convertible into a sound wave.”, and Par. 0044:”… The first user device 102 includes an example camera 330, an example display 332, an example loudspeaker 334, an example microphone 336, and the translation app 108. The second user device 104 includes an example camera 360, an example display 362, and the translation app 108. In some examples, the first user device 102 and the second user device 104 include the same components.”, and Par. 0071:” The communication server 302 transmits the audio data to the first user device 102 via the translation app 108. The translation app 108 presents the audio data to the hearing person via the loudspeaker 334 or other type of electroacoustic transducer of the first user device 102. In some examples, the communication server 302 transmits the text to the first user device 102 via the translation app 108 for presentation to the hearing person on the display 332.”, and Par. 0077:”… the translation app 108 sends the images or video captured from the camera 360 to the communication server 302.”) Regarding claims 8, and 16, Rahmani, as modified above, teaches the computer-implemented method, and the system claims of 7, and 15, respectively. Rahmani, as modified above, further teaches wherein the text prediction comprises text and a sentiment prediction and the text-to-speech model determines the audio signal from a set of defined audio signals based on the text and the sentiment prediction. (Rahmani, Par. 0069:’ The sign/text engine circuitry 314 implements the sign/text model to identify text corresponding to the sign and/or the gloss label. In addition, the sign/text engine circuitry 314 may identify non-manual features in the video of the sign language (e.g., rhythm, cadence, eye gaze, facial expressions, and/or pauses) that are analogous to intonation in spoken language. The sign/text engine circuitry 314 can tag the text with metadata indicative of one or more of the non-manual features.”, and Par. 0070:”The TTS engine circuitry 312 converts the text to audio data. In some examples, the TTS engine circuitry 312 tokenizes the text and identifies phonetic transcriptions to the words in the text. The TTS engine circuitry 312 converts the phonetic transcriptions into a sound wave (as audio data). In some examples, the TTS engine circuitry 312 uses the metadata indicative of the non-manual features of the sign language to create or adjust a volume, cadence, tone, and/or emotion of the speech.”) Regarding claim 9, Rahmani, as modified above, teaches the computer-implemented method claim of 8. Rahmani, as modified above, further teaches wherein the set of defined audio signals comprises a defined audio signal that corresponds to a text-sentiment pair associated with a corresponding textual prediction and a corresponding sentiment prediction. (Rahmani, Par. 0069:’ The sign/text engine circuitry 314 implements the sign/text model to identify text corresponding to the sign and/or the gloss label. In addition, the sign/text engine circuitry 314 may identify non-manual features in the video of the sign language (e.g., rhythm, cadence, eye gaze, facial expressions, and/or pauses) that are analogous to intonation in spoken language. The sign/text engine circuitry 314 can tag the text with metadata indicative of one or more of the non-manual features.”, and Par. 0070:”The TTS engine circuitry 312 converts the text to audio data. In some examples, the TTS engine circuitry 312 tokenizes the text and identifies phonetic transcriptions to the words in the text. The TTS engine circuitry 312 converts the phonetic transcriptions into a sound wave (as audio data). In some examples, the TTS engine circuitry 312 uses the metadata indicative of the non-manual features of the sign language to create or adjust a volume, cadence, tone, and/or emotion of the speech.”) Note: Emotional states are identified from a video's non-manual features. Claims 2-3, 11-12, and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Rahmani, and Liu, and in further view of Grishchenko et al. (herein “MediaPipe Holistic - Simultaneous Face, Hand and Pose Prediction, on Device”, 2020). Regarding claims 2, 11 and 18 Rahmani, as modified above, teaches the computer-implemented method, the system, and the non-transitory media claims of 1, 10, and 17, respectively. Rahmani, as modified above, further teaches wherein the parallel feature extraction model of the multi-stage machine learning architecture comprises a set of feature extraction models and generating the set of facial features comprises: generating, using a first feature extraction model of the set of feature extraction models, a facial expression feature based on a first image portion of the set of image portions; (Rahmani, Par. 0052:” … Such channels include manual features, such as for example, hand shape, movements, and pose as well as non-manuals features, such as for example, facial expression, mouth gestures, and movement of the head, shoulders, and torso.”, and Par. 0069:” … the sign/text engine circuitry 314 may identify non-manual features in the video of the sign language (e.g., rhythm, cadence, eye gaze, facial expressions, and/or pauses) that are analogous to intonation in spoken language.”, and Par. 0123:” … wherein the non-manual elements include at least one of a facial expression, an eye gaze, a rhythm, a cadence, eye gaze, or a pause.”) generating, using a second feature extraction model of the set of feature extraction models, an eye movement feature based on a second image portion of the set of image portions; and (Rahmani, Par. 0051:” … the key points are based on rhythm, cadence, eye gaze, and/or pauses. Any, all, and other features may be used to identify key points to be used in identifying a sign.”, and Par. 0069:” … the sign/text engine circuitry 314 may identify non-manual features in the video of the sign language (e.g., rhythm, cadence, eye gaze, facial expressions, and/or pauses) that are analogous to intonation in spoken language.”, and Par. 0123:”… wherein the non-manual elements include at least one of a facial expression, an eye gaze, a rhythm, a cadence, eye gaze, or a pause.”, and Par. 0131:” … wherein the non-manual elements include at least one of a facial expression, an eye gaze, a rhythm, a cadence, or a pause.”) generating, using a third feature extraction model of the set of feature extraction models, a lip position feature based on a third image portion of the set of image portions. (Rahmani, Par. 0051:” … the key points are based on movement of one or more of the hands, face, head, torso, mouthing, and/or finger spelling. In some examples, the key points are based on rhythm, cadence, eye gaze, and/or pauses.”, and Par. 0052:” … for example, facial expression, mouth gestures, and movement of the head, shoulders, and torso.”) Rahmani, as modified above, does not teach, Grishchenko teaches determining a set of image portions from the image; (Grishchenko, Page 3:” To streamline the identification of ROIs, a tracking approach similar to the one used for the standalone face and hand pipelines is utilized. This approach assumes that the object doesn't move significantly between frames, using an estimation from the previous frame as a guide to the object region in the current one. However, during fast movements, the tracker can lose the target, which requires the detector to re-localize it in the image. MediaPipe Holistic uses pose prediction (on every frame) as an additional ROI prior to reduce the response time of the pipeline when reacting to fast movements. This also enables the model to retain semantic consistency across the body and its parts by preventing a mixup between left and right hands or body parts of one person in the frame with another.”) Grishchenko is considered to be analogous to the claimed invention because it is in the same field of endeavor. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rahmani, as modified above, further in view of Grishchenko to determine a set of image portions from the image. Motivation to do so would isolate key areas ensure the linguistic nuances of the gesture rather than noisy backgrounds. Regarding claims 3, 12 and 19 Rahmani, as modified above, teaches the computer-implemented method, the system, and the non-transitory media claims of 2, 11, and 18, respectively. Rahmani, as modified above, does not teach, however, Liu further teaches wherein the facial expression feature, the eye movement feature, and the lip position feature are generated in parallel. (Liu, Par. 0037:”… facial expression may be a characteristic position of each of the face organs, or in combination of serial changes of the organs. The displacements among the user 1's eyebrows 40, eyes 41, ears 42, nose 43, or mouth 44 may be recognized to be the facial expressions. The facial expressions in the current examples are such as the variations of the eyebrows 40 shown in FIGS. 4A to 4C, the eyes 41 of FIGS. 5A to 5D including blinking single eye, alternately blinking eyes, or simultaneous [parallel] blinking eyes, and the variations of closing or opening mouth shown in FIGS. 6A to 6C, including combination of the expressions of opening mouth, closing mouth, and extending tongue, or the changes of shape as the user performing lip language or speaking.”, and Par. 0038:”… the illustrated facial expressions may be combined with the simultaneous motions of other facial organs. One of the facial expressions is such as closing single eye shown in FIG. 4A and FIG. 4B combined with opening mouth and closing mouth shown in FIGS. 6A through 6B.” Conclusion The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Xiaoyuan Gu (US20200005028A1) teaches in Par. 0006:” … transform continuous video streams of human sign language gestures into text, audible speech or other digital data that can be stored, manipulated and transmitted via computer.” Examiner's Note: Examiner has cited particular columns and line numbers and/or paragraph numbers in the references applied to the claims above for the convenience of the applicant. Although the specified citations are representative of the teachings of the art and are applied to specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant in preparing responses, to fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner. In the case of amending the Claimed invention, Applicant is respectfully requested to indicate the portion(s) of the specification which dictate(s) the structure relied on for proper interpretation and also to verify and ascertain the metes and bounds of the claimed invention. Any inquiry concerning this communication or earlier communications from the examiner should be directed to DARIOUSH AGAHI whose telephone number is (408)918-7689. The examiner can normally be reached Monday - Thursday and alternate Fridays, 7:30-4:30 PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bhavesh Mehta can be reached on 571-272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. DARIOUSH AGAHI, P.E. Primary Examiner /DARIOUSH AGAHI/Primary Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Dec 13, 2024
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688354
AUTOMATED NOTEBOOK COMPLETION USING SEQUENCE-TO-SEQUENCE TRANSFORMER
2y 4m to grant Granted Jul 21, 2026
Patent 12682179
RESPONSE DETERMINATION BASED ON CONTEXTUAL ATTRIBUTES AND PREVIOUS CONVERSATION CONTENT
2y 3m to grant Granted Jul 14, 2026
Patent 12664363
METHOD AND SYSTEM FOR EVALUATING NON-FICTION NARRATIVE TEXT DOCUMENTS
2y 4m to grant Granted Jun 23, 2026
Patent 12657392
EXTRACTING THEMES FROM TEXTUAL DATA
2y 6m to grant Granted Jun 16, 2026
Patent 12651597
NATURAL LANGUAGE INTERFACES
4y 0m to grant Granted Jun 09, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
99%
With Interview (+29.9%)
2y 7m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 179 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month