Prosecution Insights
Last updated: October 02, 2026
Application No. 18/366,159

CONVERSATION BASED GUIDED INSTRUCTIONS DURING A VIDEO CONFERENCE

Non-Final OA §103
Filed
Aug 07, 2023
Examiner
JONES, CARISSA ANNE
Art Unit
Tech Center
Assignee
International Business Machines Corporation
OA Round
1 (Non-Final)
77%
Grant Probability
Favorable
1-2
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
27 granted / 35 resolved
+17.1% vs TC avg
Strong +33% interview lift
Without
With
+33.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
19 currently pending
Career history
62
Total Applications
across all art units

Statute-Specific Performance

§101
2.8%
-37.2% vs TC avg
§103
80.0%
+40.0% vs TC avg
§102
11.2%
-28.8% vs TC avg
§112
3.7%
-36.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 35 resolved cases

Office Action

§103
DETAILED ACTION This action is in response to the application filed 08/07/2023. Claims 1 - 20 are pending and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1 – 3, 9, 12 – 14 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Althobaiti et al. (U.S. Pub. No. 2023/0326143, hereinafter “Althobaiti”) in view of Bergenlid et al. (U.S. Pub. No. 2019/0036856, hereinafter “Bergenlid”). Regarding Claim 1, Althobaiti teaches A method for providing conversation based guided instructions during a video conference (see Althobaiti Abstract, method, and Paragraph [0048], the remote user 134 can aid the local user 120 in the video call that provides step-by-step instructions to perform a certain procedure), the method comprising: determining a first participant in the video conference requires assistance completing a task (see Althobaiti Paragraph [0030], the local user uses a camera to capture a physical object for which the local user is requesting assistance, and Paragraph [0033], The local user 120 may then use the local device 102 to request assistance from the remote user 134); determining a second participant in the video conference is providing an instruction on how to complete the task (see Althobaiti Paragraph [0060], At step 608, method 600 involves receiving, from the remote device, data indicative of instructions for completing the task, and Paragraph [0048], The local user 120 can view the modifications in real-time on the video stream. Additionally, the remote user 134 can provide instructions to the local user 120 by drawing on the screen during the video stream view and provide written instructions); obtaining the instruction provided by the second participant (see Althobaiti Paragraph [0061], At step 610, method 600 involves modifying, based on the received data, the AR overlay displayed on the display device to provide the instructions for completing the task); and displaying a visual representation of the instruction on a display of the first participant (see Althobaiti Paragraph [0040], The local user 120 can provide instructions via a graphical user interface (GUI) displayed on the display device 106, and Paragraph [0048], Additionally, the remote user 134 can provide instructions to the local user 120 by drawing on the screen during the video stream view and provide written instructions). Althobaiti does not expressively teach determining a first participant in the video conference requires assistance completing a task on an application; However, Bergenlid teaches determining a first participant in the video conference requires assistance completing a task on an application (see Bergenlid Paragraph [0036], providing assistance during a computer-mediated communication session conducted between participant users. Providing assistance may include providing information items that are suitable in the context of a conversation between participants in the communication session, Paragraph [0037], assistant application to determine context based on media content during the communication session, e.g., audio, video, and/or text exchanged between participant users); It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a method for providing guidance during a video conference by having one participant assist another with completing a task, then displaying the provided instructions visually on the participant’s screen (as taught in Althobaiti), with determining a first participant in the video conference requires assistance completing a task on an application (as taught in Bergenlid), the motivation being to implement contextual and immediate assistance where the need arises within a video conference, rather than requiring a user to use a different application to seek assistance (see Bergenlid Paragraph [0036]). Regarding Claim 2, Althobaiti in view of Bergenlid teaches The method of claim 1, wherein the instruction includes an identified user interface element of the application being displayed on the display of the first participant (see Althobaiti Paragraph [0024], FIG. 4A and FIG. 4B illustrate an example of augmented reality indicators displayed on a graphical user interface, and Paragraph [0040], The local user 120 can provide instructions via a graphical user interface (GUI) displayed on the display device 106, and Paragraph [0048], Additionally, the remote user 134 can provide instructions to the local user 120 by drawing on the screen during the video stream view and provide written instructions). Regarding Claim 3, Althobaiti in view of Bergenlid teaches The method of claim 2, further comprising displaying an indicator on the display of the first participant adjacent to the identified user interface element (see Althobaiti Paragraph [0024], FIG. 4A and FIG. 4B illustrate an example of augmented reality indicators displayed on a graphical user interface, in which written instructions and drawings are displayed next to and around object of interest). Regarding Claim 9, Althobaiti in view of Bergenlid teaches The method of claim 1, wherein the application is a video conference application being used for the video conference (see Althobaiti Paragraph [0004], This disclosure describes methods and systems for remote assistance video communication, Paragraph [0036], the local device 102 includes an AR module 112, which can be used by the local device 102 to conduct AR video calls, e.g., with the remote device 104. As shown in FIG. 1). Regarding Claims 12 – 14, they are rejected similarly as Claims 1 – 3, respectively. The system can be found in Althobaiti (Abstract, system). Regarding Claim 20, it is rejected similarly as Claim 1. The computer program product can be found in Althobaiti (Paragraph [0092], computer program). Claims 4, 5, 15 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Althobaiti et al. (U.S. Pub. No. 2023/0326143, hereinafter “Althobaiti”) in view of Bergenlid et al. (U.S. Pub. No. 2019/0036856, hereinafter “Bergenlid”) and Shires et al. (U.S. Pub. No. 2019/0121851, hereinafter “Shires”). Regarding Claim 4, Althobaiti in view of Bergenlid teaches all the limitations of claim 1, but does not expressively teach The method of claim 1, wherein determining that the first participant requires assistance completing the task comprises identifying that the first participant is confused by performing facial expression classification on a video of the first participant. However, Shires teaches The method of claim 1, wherein determining that the first participant requires assistance completing the task comprises identifying that the first participant is confused by performing facial expression classification on a video of the first participant (see Shires Paragraph [0061], the user 102 may furl his brow while listening to another user speak (e.g., an address, a complicated series of numbers, a name with a difficult spelling), and the detector module 118 may determine that the user is concerned or confused, and provide transcript of speech to clarify speech that was spoken). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a method for providing guidance during a video conference by having one participant assist another with completing a task in an application, then displaying the provided instructions visually on the participant’s screen (as taught in Althobaiti in view of Bergenlid), with determining that a participant requires help by analyzing their facial expressions in the video to identify signs of confusion (as taught in Shires), the motivation being to implement proactive assistance, such as help being triggered when confusion is detected, rather than requiring the participant to request help (see Shires Paragraph [0061]). Regarding Claim 5, Althobaiti in view of Bergenlid and Shires teaches The method of claim 1, wherein determining that the first participant requires assistance completing the task comprises identifying that the first participant is confused by performing speech classification of a speech of the first participant (see Shires Paragraph [0066], detecting an emotion of one or more of the attending individuals. For example, a user's choice of words, patterns of speech (e.g., speed, volume, pitch, pronunciation, voice stress), gestures, facial expressions, physical characteristics, body language, and other appropriate characteristics can be detected by an emotion detection process 538 to estimate the user's emotional state). Regarding Claims 15 - 16, they are rejected similarly as Claims 4 - 5, respectively. The system can be found in Althobaiti (Abstract, system). Claims 6 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Althobaiti et al. (U.S. Pub. No. 2023/0326143, hereinafter “Althobaiti”) in view of Bergenlid et al. (U.S. Pub. No. 2019/0036856, hereinafter “Bergenlid”) and Horvitz et al. (U.S. Patent No. 6,260,035, hereinafter “Horvitz”). Regarding Claim 6, Althobaiti in view of Bergenlid teaches all the limitations of claim 1, but does not expressively teach The method of claim 1, wherein determining that the first participant requires assistance completing the task comprises determining that the first participant is searching a user interface of the application by analyzing a movement of a mouse of the first participant. However, Horvitz teaches The method of claim 1, wherein determining that the first participant requires assistance completing the task comprises determining that the first participant is searching a user interface of the application by analyzing a movement of a mouse of the first participant (see Horvitz Column 7, lines 31 – 51, The first step 50 of the process is to identify the basic functionality of a particular software program in which users experience difficulties, or could benefit from automated assistance, and the user behavior exhibited by the user when the user is experiencing those difficulties or desire for assistance. The user exhibits the following behavior: First the user selects the bar chart by placing the mouse pointer on the chart and double clicking the mouse button. Then the user pauses to dwell on the chart for some period of time while introspecting or searching for the next step. Observation of this activity could serve as an indication that the user is experiencing difficulty in using the chart feature of the spreadsheet). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a method for providing guidance during a video conference by having one participant assist another with completing a task in an application, then displaying the provided instructions visually on the participant’s screen (as taught in Althobaiti in view of Bergenlid), with determining that a participant requires assistance by analyzing the participant’s mouse movements to identify when they are searching or navigating the application’s user interface (as taught in Horvitz), the motivation being to implement a system that autonomously senses that a participant may need assistance in using a particular feature or to accomplish a specific task, and that offers to provide relevant assistance based upon considering multiple pieces of evidence (see Horvitz Column 2, lines 64 - 67). Regarding Claim 17, it is rejected similarly as Claim 6. The system can be found in Althobaiti (Abstract, system). Claims 7, 8, 10, 18 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Althobaiti et al. (U.S. Pub. No. 2023/0326143, hereinafter “Althobaiti”) in view of Bergenlid et al. (U.S. Pub. No. 2019/0036856, hereinafter “Bergenlid”) and Patel et al. (U.S. Pub. No. 2022/0376938, hereinafter “Patel”). Regarding Claim 7, Althobaiti in view of Bergenlid teaches all the limitations of claim 1, but does not expressively teach The method of claim 1, wherein determining that the second participant in the video conference is providing the instruction on how to complete the task comprises performing speech-to-text recognition of a speech of the second participant and natural language processing on the speech of the second participant. However, Patel teaches The method of claim 1, wherein determining that the second participant in the video conference is providing the instruction on how to complete the task comprises performing speech-to-text recognition of a speech of the second participant and natural language processing on the speech of the second participant (see Patel Paragraph [0015], the intelligent meeting assistant system may monitor a virtual meeting by analyzing or processing a data stream (or data packets) associated with the virtual meeting, such as by performing speech-to-text conversion processing, natural language processing, image-based processing, and/or the like on the data stream. The intelligent meeting assistant system may provide notifications of detected trigger phrase(s)/condition(s) in any suitable manner, including, for example, visually (e.g., via a pop-up message and/or by lighting up, blinking, or changing a color of an icon, a window, etc.), audibly (e.g., via an output of an audible tone, speech, etc.), haptically, and/or the like). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a method for providing guidance during a video conference by having one participant assist another with completing a task in an application, then displaying the provided instructions visually on the participant’s screen (as taught in Althobaiti in view of Bergenlid), with implementing speech-to-text recognition and natural language processing to analyze a participant’s speech (as taught in Patel), the motivation being to automate the process of interpreting speech of a participant to identify relevant information (see Patel Abstract and Paragraph [0015]). Regarding Claim 8, Althobaiti in view of Bergenlid and Patel teaches The method of claim 1, wherein obtaining the instruction provided by the second participant comprises performing a speech-to-text recognition on a speech of the second participant (see Patel Paragraph [0015], the intelligent meeting assistant system may monitor a virtual meeting by analyzing or processing a data stream (or data packets) associated with the virtual meeting, such as by performing speech-to-text conversion processing, natural language processing, image-based processing, and/or the like on the data stream. The intelligent meeting assistant system may provide notifications of detected trigger phrase(s)/condition(s) in any suitable manner, including, for example, visually (e.g., via a pop-up message and/or by lighting up, blinking, or changing a color of an icon, a window, etc.), audibly (e.g., via an output of an audible tone, speech, etc.), haptically, and/or the like). Regarding Claim 10, Althobaiti in view of Bergenlid and Patel teaches The method of claim 1, wherein the visual representation of the instruction is displayed as a pop-up window on the display of the first participant (see Patel Paragraph [0015], The intelligent meeting assistant system may provide notifications of detected trigger phrase(s)/condition(s) in any suitable manner, including, for example, visually (e.g., via a pop-up message and/or by lighting up, blinking, or changing a color of an icon, a window, etc.), audibly (e.g., via an output of an audible tone, speech, etc.), haptically, and/or the like). Regarding Claims 18 - 19, they are rejected similarly as Claims 7 - 8, respectively. The system can be found in Althobaiti (Abstract, system). Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Althobaiti et al. (U.S. Pub. No. 2023/0326143, hereinafter “Althobaiti”) in view of Bergenlid et al. (U.S. Pub. No. 2019/0036856, hereinafter “Bergenlid”) and Choi et al. (U.S. Pub. No. 2023/0245410, hereinafter “Choi”). Regarding Claim 11, Althobaiti in view of Bergenlid teaches all the limitations of claim 1, but does not expressively teach The method of claim 1, further comprising: determining that the first participant has completed the task; displaying an option for the first participant to save the textual representation of the instruction; and removing the visual representation of the instruction from the display of the first participant. However, Choi teaches The method of claim 1, further comprising: determining that the first participant has completed the task (see Choi Paragraph [0034], For instance, and in some examples, the annotation or animation may remain until the user completes the task. In alternative examples, the annotation or animation may remain until the user provides input to dismiss the annotation or animation (e.g., gesture, buttons, verbal commands, etc.). In alternative examples, the annotation or animation may remain until the customer service representative receives feedback that the task has been completed, Paragraph [0036], the annotation or animation may be cleared when the computing system receives an indication that the user completed the task corresponding to the annotation/animation, and Paragraph [0043], The computing system of the user may automatically progress through the instructions in response to detecting completion of a corresponding task); displaying an option for the first participant to save the textual representation of the instruction (see Choi Paragraph [0031], these annotations and/or animations can be saved for future reference and/or replaying or resending (e.g., if the user requires additional annotation to complete a given step). The annotations and/or animations can be saved in an archive that is accessible by a single user, or by multiple users and/or user accounts, and Paragraph [0044], may save the instructions to computing system 100 such that the user can access the instructions at a later time and repair the product (e.g., when the computing system detects the product and the tools necessary for the repair)); and removing the visual representation of the instruction from the display of the first participant (see Choi Paragraph [0040], the annotation or animation ceases to be presented). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a method for providing guidance during a video conference by having one participant assist another with completing a task in an application, then displaying the provided instructions visually on the participant’s screen (as taught in Althobaiti in view of Bergenlid), with determining that a participant has completed a task, displaying an option to save the textual representation of the instruction, and then removing the instruction from the participant’s display (as taught in Choi), the motivation being to provide the ability for the participant to save the task instruction for access to review and reuse at a later time (see Choi Paragraph [0044]). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Refer to PTO-892, Notice of References Cited for a listing of analogous art. Any inquiry concerning this communication or earlier communications from the examiner should be directed to CARISSA A JONES whose telephone number is (703)756-1677. The examiner can normally be reached Telework M-F 6:30 AM - 4:00 PM CT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached at 5712727503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CARISSA A JONES/Examiner, Art Unit 2691 /DUC NGUYEN/Supervisory Patent Examiner, Art Unit 2691
Read full office action

Prosecution Timeline

Aug 07, 2023
Application Filed
Dec 18, 2023
Response after Non-Final Action
Sep 23, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12744981
METHODS, SYSTEMS, APPARATUSES, AND DEVICES FOR FACILITATING MANAGING THE BODY LANGUAGE OF A USER DURING A VIDEO COMMUNICATION SESSION
2y 8m to grant Granted Sep 22, 2026
Patent 12737846
ITERATIVE BACKGROUND GENERATION FOR VIDEO STREAMS
3y 0m to grant Granted Sep 15, 2026
Patent 12699454
TERMINAL APPARATUS, COMMUNICATION SYSTEM
2y 11m to grant Granted Aug 04, 2026
Patent 12666140
CONTRIBUTION-BASED CLOSE-UP CONTROL
2y 3m to grant Granted Jun 23, 2026
Patent 12656991
SYSTEMS, METHODS, APPARATUSES, AND DEVICES FOR MANAGING A GAZE OF A USER ENGAGED IN A VISUAL COMMUNICATION SESSION WITH AN AUDIENCE
2y 11m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
77%
Grant Probability
99%
With Interview (+33.3%)
2y 7m (~0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 35 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month