Prosecution Insights
Last updated: August 17, 2026
Application No. 18/994,890

AUDIO AND VIDEO CALLING METHOD AND APPARATUS

Non-Final OA §103
Filed
Jan 15, 2025
Priority
Jul 15, 2022 — CN 202210840292.4 +1 more
Examiner
TSENG, CHARLES
Art Unit
Tech Center
Assignee
ZTE Corporation
OA Round
1 (Non-Final)
79%
Grant Probability
Favorable
1-2
OA Rounds
11m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
556 granted / 702 resolved
+19.2% vs TC avg
Strong +32% interview lift
Without
With
+31.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 6m
Avg Prosecution
26 currently pending
Career history
718
Total Applications
across all art units

Statute-Specific Performance

§101
13.7%
-26.3% vs TC avg
§103
53.4%
+13.4% vs TC avg
§102
6.2%
-33.8% vs TC avg
§112
13.7%
-26.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 702 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Objections Claims 8-9, 18 and 23 are objected to because of the following informalities: For claim 8, Examiner believes this claim should be amended in the following manner: The method according to claim 6, wherein before copying, by the media server, to the AI component the audio stream and the video stream of the audio/video call between the calling user and the called user, the method further comprises: receiving, by the media server, a request instruction issued by a service application for copying the audio stream and the video stream to the AI component, the request instruction carrying an audio stream identification (ID), a video stream ID, and a uniform resource locator (URL) address of the AI component; negotiating, by the media server, with the AI component port information and media information for receiving the audio stream and the video stream; and receiving, by the media server, the URL address and the port information for receiving the audio stream and the video stream, which are returned by the AI component. For claim 9, Examiner believes this claim should be amended in the following manner: The method according to claim 6, wherein superimposing, by the media server, on the audio/video call between the calling user and the called user the animation effect corresponding to the specific content comprises: receiving, by the media server, a media processing instruction from a service application, and obtaining the animation effect according to a uniform resource locator (URL) of the animation effect carried in the media processing instruction; and encoding and synthesizing, by the media server, the animation effect with the audio stream and/or the video stream, and issuing the encoded and synthesized audio stream and video stream to the calling user and the called user. For claim 18, Examiner believes this claim should be amended in the following manner: A non-transitory computer-readable storage medium, storing a computer program, wherein the computer program, when being executed by a processor, implements the method according to claim 1. For claim 23, Examiner believes this claim should be amended in the following manner: A non-transitory computer-readable storage medium, storing a computer program, wherein the computer program, when being executed by a processor, implements the method according to claim 6. Appropriate correction is required. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1, 6, 7, 18, 19, 23 and 24 is/are rejected under 35 U.S.C. 103 as being unpatentable over Todd (WO 2019/191082 A2) (made of record of the IDS submitted 10/02/2025) in view of Panton (U.S. Patent Application Publication 2016/0072778 A1). For claim 1, Todd discloses a method for audio/video calling (disclosing a method for video and voice calling (par. 398)), the method comprising: after an audio/video call between a calling user and a called user is connected to a media server, receiving, by an artificial intelligence (AI) component, an audio stream and a video stream of the audio/video call between the calling user and the called user, which are copied by the media server (disclosing, after a video conference call between users is handled by an interactive multi layer content platform as a media server (par. 427 and 584), artificial intelligence (AI) systems as an AI component receives an audio stream and a video stream of the video conference call as copied by the interactive multi layer content platform (Figs. 55A-B; par. 207, 547 and 584-585)); and recognizing, by the Al component, specific content in the audio stream and/or the video stream, and superimposing, by the Al component, on the audio/video call between the calling user and the called user an animation effect corresponding to the specific content (disclosing the AI systems recognize specific content in the audio stream and video stream (par. 574 and 584); explaining the AI systems as incorporated as part of the interactive multi layer platform overlays an animation effect corresponding to the specific content for superimposition on the video conference call (par. 234, 542, 595-596 and 607; see e.g. claim 117). Todd does not specifically disclose a call is anchored to a server. However, these limitations are well-known in the art as disclosed in Panton. Panton similarly discloses a server for handling calling between users (par. 32). Panton explains its video calls are anchored at the server (par. 32). It follows Todd may be accordingly modified with the teachings of Panton to anchor its audio/video call to its media server. A person having ordinary skill in the art (PHOSITA) before the effective filing date of the claimed invention would find it obvious to modify Todd with the teachings of Panton. Panton is analogous art in dealing with a server for handling calling between users (par. 32). Panton discloses its use of call anchoring is advantageous in appropriately managing calls and communication between users (par. 32). Consequently, a PHOSITA would incorporate the teachings of Panton into Todd for appropriately managing calls and communication between users. Therefore, claim 1 is rendered obvious to a PHOSITA before the effective filing date of the claimed invention. For claim 6, Todd as modified by Panton discloses a method for audio/video calling (Todd discloses a method for video and voice calling (par. 398)), the method comprising: after an audio/video call between a calling user and a called user is anchored to a media server, copying, by the media server, to an artificial intelligence (AI) component an audio stream and a video stream of the audio/video call between the calling user and the called user (Todd discloses, after a video conference call between users is handled by an interactive multi layer content platform as a media server (par. 427 and 584), artificial intelligence (AI) systems as an AI component receives an audio stream and a video stream of the video conference call as copied by the interactive multi layer content platform (Figs. 55A-B; par. 207, 547 and 584-585); Panton similarly discloses a server for handling calling between users (par. 32); Panton explains its video calls are anchored at the server (par. 32); and it follows Todd may be accordingly modified with the teachings of Panton to anchor its audio/video call to its media server); and according to a recognition result of specific content in the audio stream and/or the video stream by the AI component, superimposing, by the media server, on the audio/video call between the calling user and the called user an animation effect corresponding to the specific content (Todd discloses the AI systems recognize specific content in the audio stream and video stream (par. 574 and 584); Todd explains the AI systems as incorporated as part of the interactive multi layer platform overlays an animation effect corresponding to the specific content for superimposition on the video conference call (par. 234, 542, 595-596 and 607; see e.g. claim 117). For claim 7, Todd as modified by Panton discloses wherein before copying, by the media server, to the AI component the audio stream and the video stream of the audio/video call between the calling user and the called user, the method further comprises: allocating, by the media server, media resources to the calling user and the called user respectively according to an application of a call platform, such that the call platform re-anchors the calling user and the called user to the media server respectively according to the applied media resources for the calling user and the called user (Todd explains its interactive multi layer platform may be operated in accordance with an application of a call platform such as FaceTime, Skype and the like to allocate cloud locations as media resources to the users for handling the video conference call between the users (par. 398 and 599); Panton similarly discloses a server for handling calling between users (par. 32); Panton explains its video calls are anchored at the server (par. 32); and it follows Todd may be accordingly modified with the teachings of Panton to re-anchor its users to its media server according to the applied media resources). For claim 18, Todd as modified by Panton discloses a non-transitory computer-readable storage medium, storing a computer program, wherein the computer program, when being executed by a processor (Todd discloses memory for storing a computer program for execution by a processor to carry out the functions of a computer (par. 614)), implements the method according to 1 (see above as to claim 1). For claim 19, Todd as modified by Panton discloses an electronic device, comprising a memory, a processor, and a computer program stored on the memory and capable of running on the processor, wherein the computer program, when being executed by the processor (Todd discloses an electronic device including memory, a processor, and a program stored on the memory for execution by the processor to perform the functions of the electronic device (par. 614 and 622)), causes the processor to execute the following operations of claim 1 (see above as to claim 1). For claim 23, Todd as modified by Panton discloses a non-transitory computer-readable storage medium, storing a computer program, wherein the computer program, when being executed by a processor(Todd discloses memory for storing a computer program for execution by a processor to carry out the functions of a computer (par. 614)), implements the method according to 6 (see above as to claim 6). For claim 24, Todd as modified by Panton discloses an electronic device, comprising a memory, a processor, and a computer program stored on the memory and capable of running on the processor, wherein the computer program, when being executed by the processor (Todd discloses an electronic device including memory, a processor, and a program stored on the memory for execution by the processor to perform the functions of the electronic device (par. 614 and 622)), implements the method according to claim 6 (see above as to claim 6). Claim(s) 2, 20 and 25 is/are rejected under 35 U.S.C. 103 as being unpatentable over Todd view of Panton further in view of Martinez et al. (U.S. Patent Application Publication 2011/0271002 A1, hereinafter “Martinez”). For claim 2, depending on claim 1, Todd as modified by Panton does not disclose negotiation with a server port information and media information for receiving a stream, and returning a uniform resource locator (URL) address and the port information for receiving the stream. However, these limitations are well-known in the art as disclosed in Martinez. Martinez similarly discloses a system and method for managing a video conference stream between users (par. 17 and 30). Martinez explains its system implements a negotiation with a server for port information and media information for receiving the video conference stream and returning an URL address and the port information for receiving the video conference stream (par. 57, 70, 75-76 and 85). It follows Todd and Panton may be accordingly modified with the teachings of Martinez to negotiate, by its AI component, with its media server port information and media information for receiving its audio stream and its video stream, and returning, by its AI component, to the media server a URL address and the port information for receiving its audio stream and its video stream. A PHOSITA before the effective filing date of the claimed invention would find it obvious to modify Todd and Panton with the teachings of Martinez. Martinez is analogous art in dealing with a system and method for managing a video conference stream between users (par. 17 and 30). Martinez discloses its use of a negotiation request is advantageous in determining a port to facilitate an appropriate connection for receiving a video conference stream (par. 75-76). Consequently, a PHOSITA would incorporate the teachings of Martinez into Todd and Panton for determining a port to facilitate an appropriate connection for receiving a video conference stream. Therefore, claim 2 is rendered obvious to a PHOSITA before the effective filing date of the claimed invention. For claim 20, depending on claim 1, Todd as modified by Panton and Martinez discloses wherein before receiving, by the AI component, the audio stream and the video stream of the audio/video call between the calling user and the called user, which are copied by the media server, the method further comprises: receiving, by the AI component, a negotiation request from the media server (Martinez similarly discloses a system and method for managing a video conference stream between users (par. 17 and 30). Martinez explains its system implements a negotiation with a server for port information and media information for receiving the video conference stream and returning an URL address and the port information for receiving the video conference stream (par. 57, 70, 75-76 and 85). It follows Todd and Panton may be accordingly modified with the teachings of Martinez to receive, by its AI component, a negotiation request from its media server). For claim 25, depending on claim 19, this claim is a combination of the limitations of claim 19 and claim 2. It follows claim 25 is rejected for the same reasons as to claim 19 and claim 2. Claim(s) 3-5, 21 and 26-28 is/are rejected under 35 U.S.C. 103 as being unpatentable over Todd view of Panton further in view of Reponen et al. (U.S. Patent Application Publication 2008/0158334 A1, hereinafter “Reponen”) (made of record of the IDS submitted 10/02/2025). For claim 3, depending on claim 1, Todd as modified by Panton does not specifically disclose transcribing an audio into text, and sending the text to a service application, such that the service application recognizes a keyword in the text, and queries an animation effect corresponding to the keyword. However, these limitations are well-known in the art as disclosed in Reponen. Reponen similarly discloses a system and method for supplementing a video call with visualizations of visual effects such as animations (par. 7-8 and 33). Reponen explains its system implements speech recognition to transcribe audio into text where word as a keyword may be recognized in the text to query an appropriate animation effect from a table corresponding to the keyword (Fig. 4; par. 33 and 35). Reponen further explains the functions of its system may be implemented with service applications and it is understood the text may be sent to a service application for querying the appropriate animation effect from the keyword (par. 46 and 53). It follows Todd and Panton may be accordingly modified with the teachings of Reponen to transcribe, by its AI component, its audio stream into text, and sending the text to a service application for recognizing a keyword in the text and querying an animation effect corresponding to the keyword. A PHOSITA before the effective filing date of the claimed invention would find it obvious to modify Todd and Panton with the teachings of Reponen. Reponen is analogous art in dealing with a system and method for supplementing a video call with visualizations of visual effects such as animations (par. 7-8 and 33). Reponen discloses its use of speech recognition is advantageous in transcribing audio into corresponding text for determining an appropriate visual effect (par. 33 and 35). Consequently, a PHOSITA would incorporate the teachings of Reponen into Todd and Panton for transcribing audio into corresponding text for determining an appropriate visual effect. Therefore, claim 3 is rendered obvious to a PHOSITA before the effective filing date of the claimed invention. For claim 4, depending on claim 1, Todd as modified by Panton and Reponen discloses wherein recognizing, by the Al component, the specific content in the audio stream and/or the video stream further comprises: recognizing, by the Al component, a specific action in the video stream, and sending a recognition result to a service application, such that the service application queries an animation effect corresponding to the specific action (Reponen similarly discloses a system and method for supplementing a video call with visualizations of visual effects such as animations (par. 7-8 and 33); Reponen explains its system implements motion detectors to recognize gestures as specific actions where the gestures are used to query appropriate animation effects from a table (Fig. 7; par. 33 and 37); Reponen further explains the functions of its system may be implemented with service applications and it is understood the gestures as recognition results may be sent to a service application for querying the appropriate animation effects from the gestures (par. 46 and 53); and it follows Todd and Panton may be accordingly modified with the teachings of Reponen to recognize, by its AI component, a specific action in its video stream and sending a recognition result to a service application for querying an animation effect corresponding to the specific action). For claim 5, depending on claim 1, Todd as modified by Panton and Reponen discloses the animation effect comprises at least one of the following: a static image or a dynamic video (Reponen similarly discloses a system and method for supplementing a video call with visualizations of visual effects such as animations (par. 7-8 and 33); Reponen explains the animation effects may be static images or dynamic videos (par. 33); and it follows Todd and Panton may be accordingly modified with the teachings of Reponen to implement its animation effect with a static image or a dynamic video). For claim 21, depending on claim 6, Todd as modified by Panton and Reponen discloses wherein the animation effect comprises at least one of the following: a static image or a dynamic video (Reponen similarly discloses a system and method for supplementing a video call with visualizations of visual effects such as animations (par. 7-8 and 33); Reponen explains the animation effects may be static images or dynamic videos (par. 33); and it follows Todd and Panton may be accordingly modified with the teachings of Reponen to implement its animation effect with a static image or a dynamic video). For claim 26, depending on claim 19, this claim is a combination of the limitations of claim 19 and claim 3. It follows claim 26 is rejected for the same reasons as to claim 19 and claim 3. For claim 27, depending on claim 19, this claim is a combination of the limitations of claim 19 and claim 4. It follows claim 27 is rejected for the same reasons as to claim 19 and claim 4. For claim 28, depending on claim 19, this claim is a combination of the limitations of claim 19 and claim 5. It follows claim 28 is rejected for the same reasons as to claim 19 and claim 5. Claim(s) 9 and 22 is/are rejected under 35 U.S.C. 103 as being unpatentable over Todd view of Panton further in view of Whatmough (U.S. Patent Application Publication 2004/0160445 A1). For claim 9, depending on claim 1, Todd as modified by Panton discloses herein superimposing, by the media server, on the audio/video call between the calling user and the called user the animation effect corresponding to the specific content comprises: receiving, by the media server, a media processing instruction from a service application (Todd discloses its interactive multi layer platform may receive instructions for processing media from a service application (par. 589 and 620)); and encoding and synthesizing, by the media server, the animation effect with the audio stream and/or the video stream, and issuing the encoded and synthesized audio stream and video stream to the calling user and the called user (Todd discloses its interactive multi layer platform encodes and mixes to synthesize the animation effect with the audio stream and the video stream to issue the encoded and synthesized audio stream and video stream to the users (par. 541-542, 575, 578 and 595-596)). Todd as modified by Panton does not specifically disclose obtaining the animation effect according to a URL of the animation effect carried in the media processing instruction. However, these limitations are well-known in the art as disclosed in Whatmough. Whatmough similarly discloses a system and method for generating animation effects for display (par. 5). Whatmough explains its system implements commands as media processing instructions for carrying an URL as a network location for obtaining an animation effect (par. 92 and 137). It follows Todd and Panton may be accordingly modified with the teachings of Whatmough to carry an URL of its animation effect in its media processing instruction to obtain its animation effect. A PHOSITA before the effective filing date of the claimed invention would find it obvious to modify Todd and Panton with the teachings of Whatmough. Whatmough is analogous art in dealing with a system and method for generating animation effects for display (par. 5). Whatmough discloses its use of an URL is advantageous in appropriately obtaining an animation effect for presentation (par. 92 and 137). Consequently, a PHOSITA would incorporate the teachings of Whatmough into Todd and Panton for appropriately obtaining an animation effect for presentation. Therefore, claim 9 is rendered obvious to a PHOSITA before the effective filing date of the claimed invention. For claim 22, depending on claim 9, Todd as modified by Panton and Whatmough discloses wherein the media processing instruction is generated according to the recognition result of the specific content in the audio stream and/or the video stream by the AI component (Todd discloses its interactive multi layer platform generates instructions for processing media according to the recognition of the specific content in the audio stream or video stream by the AI systems (par. 584 and 589)). Allowable Subject Matter Claim 8 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims and to address any claim objections discussed above in the Detailed Action. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHARLES TSENG whose telephone number is (571)270-3857. The examiner can normally be reached 8-5. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached at (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CHARLES TSENG/ Primary Examiner, Art Unit 2613
Read full office action

Prosecution Timeline

Jan 15, 2025
Application Filed
Jul 29, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705850
METHOD OF RENDERING DYNAMIC LABELS IN AN EXTENDED REALITY ENVIRONMENT
11m to grant Granted Aug 11, 2026
Patent 12682582
SYSTEMS AND METHODS FOR PASS-THROUGH EXTENDED REALITY (XR) CONTENT
2y 5m to grant Granted Jul 14, 2026
Patent 12682325
SYSTEM AND METHOD FOR AUGMENTED REALITY EFFECT CUSTOMIZATION IN SOCIAL MEDIA PLATFORMS THROUGH USER-UPLOADED MEDIA
2y 4m to grant Granted Jul 14, 2026
Patent 12682525
METHOD FOR PROVIDING AVATAR AND ELECTRONIC DEVICE SUPPORTING THE SAME
2y 1m to grant Granted Jul 14, 2026
Patent 12682551
DATA PROCESSING METHOD AND SERVER BASED ON VOXEL DATA, MEDIUM AND COMPUTER PROGRAM PRODUCT
2y 3m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
79%
Grant Probability
99%
With Interview (+31.5%)
2y 6m (~11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 702 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month