Prosecution Insights
Last updated: August 18, 2026
Application No. 18/066,365

INTELLIGENT CAPTION EDGE COMPUTING

Non-Final OA §103
Filed
Dec 15, 2022
Examiner
LEE, JANGWOEN
Art Unit
2656
Tech Center
2600 — Communications
Assignee
International Business Machines Corporation
OA Round
2 (Non-Final)
85%
Grant Probability
Favorable
2-3
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 85% — above average
85%
Career Allowance Rate
45 granted / 53 resolved
+22.9% vs TC avg
Strong +18% interview lift
Without
With
+18.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
15 currently pending
Career history
75
Total Applications
across all art units

Statute-Specific Performance

§101
24.6%
-15.4% vs TC avg
§103
61.1%
+21.1% vs TC avg
§102
9.9%
-30.1% vs TC avg
§112
3.0%
-37.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 53 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments This communication is in response to the Application filed on 04/23/2026. Claims 1-20 are pending and have been examined. Claims 1, 8, and 15 are independent. Applicant’s arguments, with respect to Claims 1-20 have been fully considered and are persuasive. The rejections to Claims 1-20 under 35 U.S.C. § 103 in the previous office action have been withdrawn. Applicant's persuasive arguments necessitated the new grounds of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE NON-FINAL. Specification The disclosure is objected to because of the following informalities: In paragraph [0040], the last sentence is incomplete: “The context-based analysis may include, among other things, _.” Considering the importance of “context-based analysis,” appropriate correction is required. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-4, 7-11 and 14-18 are rejected under 35 U.S.C. 103 as being unpatentable over Maegawa (US Pub 2009/0070102) in view of Kim et al. (US Pub 2021/0358502) further in view of Manepalli et al. ("Dyn-ASR: Compact, multilingual speech recognition via spoken language and accent identification." 2021 IEEE 7th World Forum on Internet of Things (WF-IoT). IEEE, 2021). Regarding Claim 1, Maegawa discloses a method of real-time caption generation (Fig.1, par [033], "…The minutes recording unit 50 includes a minutes recording means 501 and a minutes displaying means 502…."; Fig.8, par [032], "…the speech recognition result is transmitted to the minute recording server 831 via the network 840 and is registered as the conference minutes data...") in an edge computing environment (Fig.8, par [029], "…The distribution terminal (850, ... ) is a terminal used to distribute a recognition model. Plural distribution terminals (850, ... ) are connected to the network 40..."), executable by a processor, comprising: monitoring contexts related to one or more participants in a web conference service, the one or more participants using a caption service (par [029], "…The characteristic information includes subject information, language information, relevant information relating to input speech...The language information includes dialect information..."; Figs.4 and 6, par [038], "…The model selection unit 20 searches and extracts a language information (4202, 4205 in FIG. 4) when the language information exists in the user information as a result of reading the user information ( 4201 in FIG. 4). Fig.6 shows an example that a user information table 6202 is prepared together with the conferencing data 6201..."); determining personal characteristics associated with each of the participants based on the monitored contexts (par [041], "…The model selection unit 20 mines a text of the subject information from the conferencing data 6201 ( 4301 in FIG. 4). When the language model can be identified based on the result of mining ( 4305 in FIG. 4), model selection unit 20 determines the model connection destination information in this timing…"); deploying the customized lightweight user accent-oriented caption edge module to the selected edge device based on a corpus of captions associated with the user accent-oriented caption edge module most closely matching the determined personal characteristics (Fig.1, par [033], "…The distribution server 30 includes a language model 301, an acoustic model 302, and a model distribution means 303..."; Fig.3, par [039], "…the distribution server 30 searches the relevant language based on the language information transmitted...When the servers are found, the distribution server 30 determines the acoustic model matched with the dialect information included in the language information in the process of the secondary search 361 of acoustic model and prepares 362 for distribution of the acoustic model..."). Maegawa discloses the speech recognition and a minute displaying means (i.e., caption generation) and automatic selection and distribution of the matching language/acoustic model to each participant's terminals by the conference server. Maegawa does not explicitly teaches an edge computing environment and "selecting an edge device from among a plurality of edge devices based on the determined personal characteristics, the selected edge device being configured to perform caption conversion for a participant from among the one or more participants." Kim, in the analogous field of endeavor, discloses selecting an edge device from among a plurality of edge devices based on the determined personal characteristics, the selected edge device being configured to perform caption conversion for a participant from among the one or more participants in the web conference service (Fig.4, par [008, 130-136], "…A speech recognition method in an edge computing device performed in a distributed network system including a client device, the edge computing device, and a cloud server..."; par [019], "…generating the NLU model according to a training result; compressing the NLU model to a predetermined size according to the voice pattern of the user; and transmitting the compressed NLU model to the edge computing device..."); Therefore, it would have been obvious to one of ordinary skill in the art, before effective filing date of the claimed invention, to have modified the speech recognition and a minute displaying means in conference server/terminal framework of Maegawa with the edge computing device with a user-personalized and compressed NLU model executing the speech recognition near the user of Kim with a reasonable expectation of success to directly perform a speech recognition operation with the client without involvement of the cloud. As a result, the occurrence of network overhead (latency) may be minimized (Kim, Abstract, paras [002-003, 134]). Maegawa in view of Kim does not explicitly discloses the limitation, "customizing a lightweight user accent-oriented caption edge module associated with the selected edge device for the participant." Manepalli, in the analogous field of endeavor, discloses customizing a lightweight user accent-oriented caption edge module associated with the selected edge device for the participant (Fig. 1, Abstract, 3 Approach, "…enable multilingual speech recognition on edge devices. This approach uses both language identification and accent identification to select one of multiple monolingual ASR models on-the-fly, each fine-tuned for a particular accent..."; 4.3. Fine-tuned ASR, "…The ASR models we fine-tuned were based on the QuartzNet architecture and fine-tuned on Indian, Chinese, and Malaysian accented English data from the Singapore National Speech Corpus…"); and Therefore, it would have been obvious to one of ordinary skill in the art, before effective filing date of the claimed invention, to have modified the edge computing with a personalized NLU/NLP module for conference participants of Maegawa in view of Kim with the dynamic accent-specific monolingual ASR model-selection mechanism of Manepalli with a reasonable expectation of success to reduce the need for higher end compute and memory, preserve/improve the recognition performance of monolingual ASR models, and ultimately enable multilingual contactless interactions on edge devices (Manepalli, 1. Introduction) Regarding Claim 2, Maegawa in view of Kim further in view of Manepalli discloses the method of claim 1, further comprising providing a caption conversion service from the deployed edge device to the participant from among the one or more participants (Maegawa, Fig.1, par [029], "… The distribution terminal (850, ... ) is a terminal used to distribute a recognition model. Plural distribution terminals (850, ... ) are connected to the network 40..."; Kim, Fig.4: A cloud/edge computing based adaptive speech recognition method). Regarding Claim 3, Maegawa in view of Kim further in view of Manepalli discloses the method of claim 2, further comprising causing the provided caption conversion service to display captions on one or more devices associated with the participant from among the one or more participants (Maegawa, Fig.5, par [034], "…The minutes displaying means 502 displays the minutes data on the displaying screens of users…"; Fig.8: terminals). Regarding Claim 4, Maegawa in view of Kim further in view of Manepalli discloses the method of claim 1, further comprising defining a framework to maintain and customize edge generation based on dynamically monitoring the contexts (Kim, par [241], "…the voice data of the user transferred to the edge device may be periodically monitored. The edge device may request the cloud to update the NLU model if the NLU model cannot be applied to the input voice command but the voice command is repeatedly used periodically as a result of monitoring..."). Regarding Claim 7, Maegawa in view of Kim further in view of Manepalli discloses the method of claim 1, wherein the personal characteristics correspond to one or more from among location, native language, secondary language (Maegawa, Figs. 5-7:location, site, par [029], "…The characteristic information includes subject information, language information, relevant information relating to input speech, and language information for translation. The language information includes dialect information..."; Manepalli, Abstract, "...This approach uses both language identification and accent identification to select one of multiple monolingual ASR models on-the-fly, each fine-tuned for a particular accent... ") Claim 8 is a system claim with limitations similar to the limitations of Claim 1 and is rejected under similar rationale. Additionally, Kim discloses a computer system for real-time caption generation in an edge computing environment, the computer system comprising: one or more computer-readable non-transitory storage media configured to store computer program code; and one or more computer processors configured to access said computer program code and operate as instructed by said computer program code, said computer program code including (Kim, Fig.1, par [024], "…an edge computing device includes: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include an instruction for performing the speech recognition method in the edge computing device...") … Rationale for combination is similar to that provided for Claim 1. Claim 9 is a system claim with limitations similar to the limitations of Claim 2 and is rejected under similar rationale. Claim 10 is a system claim with limitations similar to the limitations of Claim 3 and is rejected under similar rationale. Claim 11 is a system claim with limitations similar to the limitations of Claim 4 and is rejected under similar rationale. Claim 14 is a system claim with limitations similar to the limitations of Claim 7 and is rejected under similar rationale. Claim 15 is a non-transitory computer readable medium claim with limitations similar to the limitations of Claim 1 and is rejected under similar rationale. Additionally, Kim discloses a non-transitory computer readable medium having stored there on a computer program for real-time caption generation in an edge computing environment, the computer program configured to cause one or more computer processors to (Kim, Fig.1, par [026], "…A recording medium for performing an edge computing operation, as a non-transitory computer-readable medium storing a computer-executable component configured to be executed in one or more processors of a computing device...") … Rationale for combination is similar to that provided for Claim 1. Claim 16 is a non-transitory computer readable medium claim with limitations similar to the limitations of Claim 2 and is rejected under similar rationale. Claim 17 is a non-transitory computer readable medium claim with limitations similar to the limitations of Claim 3 and is rejected under similar rationale. Claim 18 is a non-transitory computer readable medium claim with limitations similar to the limitations of Claim 4 and is rejected under similar rationale. Claims 5, 12 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Maegawa in view of Kim further in view of Manepalli further in view of Streat (US Pat US10629192). Regarding Claim 5, Maegawa in view of Kim further in view of Manepalli discloses the method of claim 4. Streat, in the analogous field of the accuracy and efficiency of machine speech recognition, discloses defining a data structure for saving and tracking the caption service associated with each participant (Streat, Fig.4, col.7:1-col.8:24, "… The system 400 includes speech sample 301, speech information 401 about a user, a first data store 403 of default phoneme mappings 405, and a second data store 407 including the phoneme mapping 409 determined for the user...The speech information 401 can include data about where the user is currently located, where the user lived, the user's age, the user's gender, and the like. A default accent can be determined for the user based at least in part on the speech information 401."; Fig.5, col.11:13-15, "…At block 525, the mutated phonemes can be saved as a phoneme mapping associated with a user profile (Fig.5)..."). Therefore, it would have been obvious to one of ordinary skill in the art, before effective filing date of the claimed invention, to have modified a dynamic accent-edge speech recognition system of Maegawa in view of Kim further in view of Manepalli with the automatic characteristic inference and persistent user profile mechanism of Streat with a reasonable expectation of success to enhance the accuracy, convenience, and session-to-session consistency of the combined system's participant, accent-adapted edge system without requiring heady processing power (Streat, col.1:9-20). Claim 12 is a system claim with limitations similar to the limitations of Claim 5 and is rejected under similar rationale. Rationale for combination is similar to that provided for Claim 5. Claim 19 is a non-transitory computer readable medium claim with limitations similar to the limitations of Claim 5 and is rejected under similar rationale. Rationale for combination is similar to that provided for Claim 5. Claims 6, 13 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Maegawa in view of Kim further in view of Manepalli further in view of Sharma et al. (US Pub 2020/0327371). Regarding Claim 6, Maegawa in view of Kim further in view of Manepalli discloses the method of claim 4. Sharma, in the analogous field of intelligent edge computing for processing and analysis in distributed network IoT environments in real-time, discloses maintaining and updating user caption conversion service profiles according to the determined personal characteristics (Sharma, Fig.9, par [011], "…the system and method provide a closed loop arrangement for continuously evaluating the accuracy of the model on the edge-computing platform, generating an updated or modified model, and iteratively updating or replacing the model on the edge computing platform to improve accuracy..."; paras [011, 237], "…The edge - based model is updated without interrupting the processing of sensor stream data by defining the model as a continuous stream that flows with the sensor data streams..."). Therefore, it would have been obvious to one of ordinary skill in the art, before effective filing date of the claimed invention, to have modified a dynamic accent-edge speech recognition system of Maegawa in view of Kim further in view of Manepalli with the closed-loop, model evaluation and update framework of Sharma with a reasonable expectation of success to continuously monitor the accuracy of the edge system and update the model in order to maintain the edge devices remain accurate over time (Sharma, paras [002-010]). Claim 13 is a system claim with limitations similar to the limitations of Claim 6 and is rejected under similar rationale. Rationale for combination is similar to that provided for Claim 6. Claim 20 is a non-transitory computer readable medium claim with limitations similar to the limitations of Claim 6 and is rejected under similar rationale. Rationale for combination is similar to that provided for Claim 6. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Belisario et al. (US Pat 9552810) discloses a computer implemented method for customizing speech recognition for users with language accents includes identifying a spoken language of a user. The method includes receiving an indicator of a speech accent language initiated by the user using a computer. An automatic speech recognition (ASR) conversion is adjusted based on the speech recognition characteristics (Belisario, Summary). Any inquiry concerning this communication or earlier communications from the examiner should be directed to JANGWOEN LEE whose telephone number is (703)756-5597. The examiner can normally be reached Monday-Friday 8:00 am - 5:00 pm ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, BHAVESH MEHTA can be reached at (571)272-7453. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JANGWOEN LEE/ Examiner, Art Unit 2656 /BHAVESH M MEHTA/ Supervisory Patent Examiner, Art Unit 2656
Read full office action

Prosecution Timeline

Dec 15, 2022
Application Filed
Nov 08, 2023
Response after Non-Final Action
Jan 27, 2026
Non-Final Rejection mailed — §103
Apr 06, 2026
Interview Requested
Apr 20, 2026
Examiner Interview Summary
Apr 20, 2026
Applicant Interview (Telephonic)
Apr 23, 2026
Response Filed
Jul 17, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12706079
ARTIFICIAL INTELLIGENCE-BASED AUDIO SIGNAL GENERATION METHOD AND APPARATUS, DEVICE, AND STORAGE MEDIUM
3y 8m to grant Granted Aug 11, 2026
Patent 12706080
END-TO-END NATURAL AND CONTROLLABLE EMOTIONAL SPEECH SYNTHESIS METHODS
3y 1m to grant Granted Aug 11, 2026
Patent 12688363
SYSTEMS AND METHODS FOR IDENTIFYING AND ANALYZING RISK EVENTS FROM DATA SOURCES
2y 5m to grant Granted Jul 21, 2026
Patent 12688359
ADAPTIVE CODE CONSTRUCT GENERATION FOR DETECTING IDENTIFIERS IN MESSAGES
2y 9m to grant Granted Jul 21, 2026
Patent 12645876
AUTO-CORRECTING FRAMEWORK FOR OPEN INFORMATION EXTRACTION SYSTEMS
3y 1m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
85%
Grant Probability
99%
With Interview (+18.0%)
2y 8m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 53 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month