Prosecution Insights
Last updated: October 01, 2026
Application No. 18/881,995

VIDEO CALL CONTROL METHOD, COMMUNICATION DEVICE AND STORAGE MEDIUM

Non-Final OA §103
Filed
Jan 07, 2025
Priority
Jul 08, 2022 — CN 202210800411.3 +1 more
Examiner
NGUYEN, VIET
Art Unit
Tech Center
Assignee
ZTE Corporation
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
3 currently pending
Career history
6
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Interpretation 2. The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: "virtual human service unit" in claim 1-2, 5, 7-9, 14-15, and 18-20. “call session control function unit” in claim 1, 5, 7, and 15. “session access unit” in claim 1, 5-8, 11, 15-16, and 18. Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. Claim 1: “a virtual human service unit” corresponds to Fig. 8 (See Fig. 8 Below, PNG media_image1.png 270 396 media_image1.png Greyscale ) Claims 2, 5, 7-9, 14-15, and 18-20 inherently invoke 112f due to their dependency from claim 1. Claim 1: “call session control function unit” corresponds to ([0032] “…a first Inquiry/Service Call Session Control Function (I/S-CSCF) network element…) Claims 5, 7, and 15 inherently invoke 112f due to their dependency from Claim 1. Claim 1: “session access unit” corresponds to ([0024] “…and sending the virtual human video data to a unit, so as to send the virtual human video data to a second terminal device via the unit.”). Claims 1, 5-8, 11, 15-16, and 18 inherently invoke 112f due to their dependency from Claim 1. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Shan et al (CN 109740476) in view of Yang et al (WO 2022062625). Consider Claim 1. Shan teaches acquiring a video call connection request of a first terminal device sent by the call session control function unit (Pg. 3 line 42-43, “after the first user connects to the second user through a user terminal including mobile phone applications, computer software, etc., after the video call is made, the mobile phone”, e.g. video call connection request included in a vide call, phone interpreted as a first terminal device, call session control unit interpreted as an application or software, Shan); sending the video call connection request to the virtual human service unit (Pg. 4 line 11, “in the function menu of the mobile phone application or computer software.”, e.g. function menu of application or software interpreted as a virtual human service unit”, Shan), and acquiring a virtual human image corresponding to the first terminal device (Pg. 1 line 22, “generate a virtual portrait model according to the human image of the first user”, e.g. generate virtual portrait mode interpreted as acquiring a virtual human image, e.g. first user is interpreted as corresponding to the first terminal device , Shan) sent by the virtual human service unit; performing a replacement processing on a target object (Pg. 4 line 8, “In the menu, the character image of the first user in the video call is replaced with a virtual image, which can be a full-body image replacement or only an avatar replacement.“, e.g. character image of first user replaced with virtual image interpreted as performing replacement processing on target object, Shan) in a video based on the virtual human image to obtain virtual human video data (Pg. 9 line 50, “generate a virtual character skeleton according to the virtual demand information”, e.g. act of generating virtual character includes virtual human video data thus interpreted as obtaining virtual human video data, Shan); and sending the virtual human video data to the session access unit, so as to send the virtual human video data to a second terminal device via the session access unit (Pg. 5 line 12, “Specifically, the server generates an enhanced video image and sends it to the second user terminal, where the second user can see the augmented reality video during the video call with the first user on the second user terminal”, e.g. server sending the enhanced video image to second user terminal interpreted as sending virtual human video data to session access unit, e.g. second user receiving the enhanced vide interpreted as the second terminal device via session access unit, Shan). Shan does not teach the video call control method specifically performed by a VoNR. Yang teaches the Voice over New Radio (VoNR+) platform (Pg. 8 lines 31-34, “The video call can be, but is not limited to, use a voice over long-term evolution (Voice over Long-Term Evolution, VoLTE) or a social application (application, APP), and can also be used for a voice over new radio (Voice over New Radio, VoNR) video call.”, e.g. video call using VoLTE can also be used for VoNR video call reads on VoNR, Yang). Although Shan teaches a conventional video call control method, Yang incorporates similar features for a VoNR platform (See Pg. 10 lines 28-30, “The CSCF may include one of a serving-call session control function (S-CSCF)…”, e.g. CSCF as call session control function, etc.) thus both references are within the same field of invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings of Yang into the teachings of Shan, to yield the result of the process behind generating and obtaining a virtual image in a video call session through a VoNR+ platform. Consider Claim 2. Shan teaches The video call control method of claim 1, wherein performing a replacement processing on a target object in a video according to the virtual human image to obtain virtual human video data comprises: acquiring video setting requirement information (Pg. 2 line 23, “where the virtual demand information is used to describe Image modification requirement”, e.g. image modification requirement interpreted as a video setting requirement information, Shan) of the first terminal device (Pg. 3 line 42-43, “after the first user connects to the second user through a user terminal including mobile phone applications, computer software, etc.”, e.g. first user connection via user terminal interpreted as first terminal device (connection requires a device), Shan) via the virtual human service unit (Pg. 5 line 38, “The video processing unit is configured to generate an enhanced video image”, e.g. video processing unit interpreted as virtual human service unit, Shan); and performing the replacement processing on the target object in the video according to the video setting requirement information and the virtual human image to obtain the virtual human video data (Pg. 2 line 23-25, “the model generation unit is configured to generate a virtual portrait model based on the personal image of the first user and the virtual requirement information; the video processing unit is configured to generate an enhanced video image, wherein The enhanced video image is obtained”, e.g. generating virtual portrait based on personal image interpreted as performing the replacement processing on target object, e.g. virtual requirement information interpreted as according to the video setting requirement, e.g. enhanced video image being obtained interpreted as obtaining virtual human video data, Shan). Consider Claim 3. Shan teaches the video call control method of claim 2, wherein performing the replacement processing on the target object in the video according to the video setting requirement information and the virtual human image to obtain the virtual human video data comprises: determining the target object to be replaced in the video according to the video setting requirement information (Pg. 3 line 4, “In the embodiment of this application, the method of extracting the human body image from the video image is adopted, and then the virtual portrait model is generated according to the human image and virtual requirements.”, e.g. virtual portrait model generated according to human image and virtual requirements is interpreted as determining target object to be replaced according to video setting requirement info., Shan), wherein the target object is a background and/or a head image in the video (Pg. 5 lines 12-13, “…which includes the real background of the first user Augmented reality video of the avatar in the image”, e.g. real background of first user interpreted as target object as background/or head image in video, Shan); and performing the replacement processing on the target object according to the virtual human image to obtain the virtual human video data (Pg. 4 lines 6-8, “In the menu, the character image of the first user in the video call is replaced with a virtual image, which can be a full-body image replacement or only an avatar replacement”, e.g. replacing character image of first user in video call with virtual image interpreted as performing replacement processing on target object, e.g. full-body image or avatar replacement interpreted as obtaining virtual human video data, Shan) Consider Claim 4. Shan teaches the video call control method of claim 2, wherein performing the replacement processing on the target object in the video according to the video setting requirement information and the virtual human image to obtain the virtual human video data comprises: according to the video setting requirement information (Pg. 2 line 23, “where the virtual demand information is used to describe Image modification requirement”, e.g. image modification requirement interpreted as a video setting requirement information, Shan), performing the replacement processing on a head image in the video with the virtual human image (Pg. 4 lines 6-8, “In the menu, the character image of the first user in the video call is replaced with a virtual image, which can be a full-body image replacement or only an avatar replacement”, e.g. replacing character image of first user with a virtual image may also include processing the head image thus interpreted as replacement processing on a head image, Shan), and performing a blurring processing on a background in the video (Pg. 8 lines 22-24, “server generates an enhanced video image and sends it to the second user terminal, where the second user can see the augmented reality video during the video call with the first user on the second user terminal, which includes the real background of the…”, e.g. generating enhanced video image containing background is interpreted as performing a blurring processing on a background, Shan), to obtain the virtual human video data. Consider Claim 5. Yang teaches the video call control method of claim 1, wherein the VoNR+ platform comprises a service module, a call capability module and a media module, and the method comprises: acquiring, by the call capability module (Pg. 10 lines 28-29, “a proxy-CSCF (Proxy CSCF, P-CSCF), an interrogating-CSCF (Interrogating-CSCF, I-CSCF), or multiple “, e.g. I-CSCF reads on call capability module, Yang), the video call connection request (Pg. 24 line 23-24, “during the process of establishing a call connection with the second terminal device”, Yang) of the first terminal device via the call session control function unit (Pg. 10 lines 33-34, “…and can provide access control, quality of service control…”, e.g. access control interpreted as call session control function unit, Yang), and sending, by the call capability module, the video call connection request to the virtual human service unit via the service module (Pg. 10 lines 12-13, “The application service function is used to interact with the AR media server”, e.g. application service function interpreted as BOTH virtual human service unit/service module, Yang); sending, by the service module, the virtual human image corresponding to the first terminal device acquired from the virtual human service unit to the media module (Pg. 10 lines 1-5, “the application server receives information from the first terminal device (such as call interface operation instructions). ), the application server sends the received information to the AR media server; thus, the AR media server processes the upstream media stream according to the information from the first terminal device”, e.g. AR media server interpreted as media module, Yang); and performing, by the media module, the replacement processing on the target object in the video with the virtual human image to obtain the virtual human video data (Pg. 8 lines 53-56, “The user video may include a video collected by a camera (a front-facing camera or a rear-facing camera) of a terminal device adopted by the user, or a file opened by the user, or an avatar of the user, and the like. For example, AR media enabler has strong image processing functions and data calculation functions, and can use AR technology to perform logical operations, screen rendering, virtual scene synthesis and other operations on the received media streams.”, e.g. screen rendering interpreted as replacement processing, e.g. avatar of user interpreted as virtual human image/virtual human video data, Yang), and sending, by the media module, the virtual human video data to the second terminal device via the session access unit (See Pg. 2, SUMMARY OF THE INVENTION lines 18-19 “A video image, superimposing the user-drawn pattern on the user's video image to obtain a third video stream, and sending the third video stream to the second terminal device”, e.g. user’s video image sent to second terminal device interpreted as sending virtual human video data to the second terminal device, Yang). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings of Yang into the teachings of Shan, to yield the result of the process behind generating and obtaining a virtual image in a video call session through a VoNR+ platform. Consider Claim 6. Yang teaches the video call control method of claim 1, wherein sending the virtual human video data to the session access unit, so as to send the virtual human video data to a second terminal device via the session access unit comprises: in response to the second terminal device being a terminal device which supports a data channel (“the terminal device may establish different auxiliary transmission channels according to the type of the auxiliary media stream to be transmitted”, e.g. terminal device interpreted as second terminal device being a terminal device, e.g. aux transmission channel interpreted as a data channel, Yang), sending the virtual human video data (Pg. 8 lines 53-56, “The user video may include a video collected by a camera (a front-facing camera or a rear-facing camera) of a terminal device adopted by the user, or a file opened by the user, or an avatar of the user, and the like.”, e.g. avatar of user interpreted as virtual human image/virtual human video data, Yang) to the session access unit (Pg. 10 lines 33-34, “…and can provide access control, quality of service control…”, e.g. access control interpreted as call session control function unit, Yang), so as to send the virtual human video data to the second terminal device through a data transmission channel (“the terminal device may establish different auxiliary transmission channels according to the type of the auxiliary media stream to be transmitted”, e.g. terminal device interpreted as second terminal device being a terminal device, e.g. aux transmission channel interpreted as a data channel, Yang) via the session access unit (Pg. 10 lines 33-34, “…and can provide access control, quality of service control…”, e.g. access control interpreted as call session control function unit, Yang); or, in response to the second terminal device being a terminal device which does not support data channel (Pg. 11 lines 19-20, “The terminal devices involved in the embodiments of this application, such as the first terminal device and the third terminal device, all support video calls, for example, support VoLTE”, e.g. second terminal device does not support video calls compared to first and third terminal device and video calls involves a data channel, Yang), sending the virtual human video data to the session access unit to perform a format conversion processing on the virtual human video data (Pg. 27 lines 2-6, “As an example, the AME1 may encapsulate the superimposed video data into data packets to obtain the superimposed video stream, and the data packets obtained by encapsulating the superimposed video data may be in RTP or RTCP format”, e.g. superimposed video data interpreted as virtual human video data, e.g. encapsulating superimposed vide data in RTP or RTCP format interpreted as performing format conversion processing on the data, e.g. AME1 interpreted as a session access unit, Yang) via the session access unit, and sending the virtual human video data subjected to the format conversion processing to the second terminal device through a video transmission channel (Pg. 27 lines 14-15 “910, The first terminal device displays the user-drawn pattern to be displayed. 911,The second terminal device displays the user-drawn pattern to be displayed and plays user audio data included in the audio data packet.”, e.g. displaying user-drawn pattern and user audio involves video transmission channel, Yang). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings of Yang into the teachings of Shan, to yield the result of the process behind generating and obtaining a virtual image in a video call session through a VoNR+ platform. Consider Claim 7. Shan teaches acquiring video data corresponding to the second terminal device sent by the session access unit (Pg. 5 line 12, “Specifically, the server generates an enhanced video image and sends it to the second user terminal, where the second user can see the augmented reality video during the video call with the first user on the second user terminal”, e.g. server sending the enhanced video image to second user terminal interpreted as acquiring/sending virtual human video data to session access unit, e.g. second user receiving the enhanced video interpreted as the second terminal device via session access unit, Shan). Shan does not teach a Voice over New Radio (VoNR+) platform, a virtual human service unit, a call session control function unit and a session access unit , wherein the virtual human service unit, the call session control function unit and the session access unit are connected to the VoNR+ platform, the method comprising: sending, to the session access unit, a video call connection request for a video call with a second terminal device, in turn sending the video call connection request to the virtual human service unit sequentially via the session access unit, the call session control function unit, and the VoNR+ platform, so that the VoNR+ platform acquires a virtual human image corresponding to the first terminal device via the virtual human service unit, performs a replacement processing on a target object in a video based on the virtual human image to obtain virtual human video data to obtain virtual human video data, and sends the virtual human video data to the second terminal device via the session access unit; and acquiring video data corresponding to the second terminal device sent by the session access unit. Yang teaches a video call control method, performed by a first terminal device connected to a communication network comprising a Voice over New Radio (VoNR+) platform (Pg. 8 lines 31-33, “The video call can be, but is not limited to, use a voice over long-term evolution (Voice over Long-Term Evolution, VoLTE) or a social application (application, APP), and can also be used for a voice over new radio (Voice over New Radio, VoNR) video call.”, Yang), a virtual human service unit, a call session control function unit and a session access unit , wherein the virtual human service unit, the call session control function unit and the session access unit are connected to the VoNR+ platform, the method comprising: sending, to the session access unit (Pg. 10 lines 27-28, “…implements functions such as user access, authentication, session routing…”, e.g. user access interpreted as session access unit, Yang), a video call connection request for a video call with a second terminal device (Pg. 24 line 23-24, “during the process of establishing a call connection with the second terminal device”, Yang), in turn sending the video call connection request to the virtual human service unit (Pg. 10 lines 12-13, “The application service function is used to interact with the AR media server”, e.g. application service function interpreted as BOTH virtual human service unit/service module, Yang) sequentially via the session access unit (Pg. 10 lines 27-28, “…implements functions such as user access, authentication, session routing…”, e.g. user access interpreted as session access unit, Yang), the call session control function unit (Pg. 10 lines 33-34, “…and can provide access control, quality of service control…”, e.g. access control interpreted as call session control function unit, Yang), and the VoNR+ platform, so that the VoNR+ platform acquires a virtual human image (Pg. 10 lines 45-46, “The virtual model, for example, may include one or more of a virtual portrait model, and material images…”, e.g. virtual portrait model interpreted as virtual human image, Yang) corresponding to the first terminal device (Pg. 10 lines 1-5, “the application server receives information from the first terminal device…”, Yang) via the virtual human service unit (Pg. 10 lines 12-13, “The application service function is used to interact with the AR media server”, e.g. application service function interpreted as virtual human service unit module, Yang), performs a replacement processing on a target object in a video based on the virtual human image to obtain virtual human video data (Pg. 8 lines 53-56, “The user video may include a video collected by a camera (a front-facing camera or a rear-facing camera) of a terminal device adopted by the user, or a file opened by the user, or an avatar of the user, and the like. For example, AR media enabler has strong image processing functions and data calculation functions, and can use AR technology to perform logical operations, screen rendering, virtual scene synthesis and other operations on the received media streams.”, e.g. screen rendering interpreted as replacement processing, e.g. avatar of user interpreted as virtual human image/virtual human video data, Yang), and sends the virtual human video data to the second terminal device via the session access unit (See Pg. 2, SUMMARY OF THE INVENTION lines 18-19 “A video image, superimposing the user-drawn pattern on the user's video image to obtain a third video stream, and sending the third video stream to the second terminal device”, e.g. user’s video image sent to second terminal device interpreted as sending virtual human video data to the second terminal device, Yang). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings of Yang into the teachings of Shan, to yield the result of the process behind generating and obtaining a virtual image in a video call session through a VoNR+ platform. Consider Claim 8. Yang teaches the video call control method of claim 7, wherein before sending, to the session access unit, a video call connection request for a video call with a second terminal device, the method further comprises: sending the virtual human image (Pg. 10 lines 40-42, “The auxiliary media stream may also include one or more of point cloud data, spatial data (which may also be referred to as spatial pose data), user-view video, or a virtual model”, e.g. virtual mode interpreted as virtual human image, Yang) and number information of the first terminal device to the virtual human service unit (Pg. 5 lines 44-45, “the first terminal device sends the first data stream and the second data stream to the media server”, e.g. first data stream sent from first terminal device interpreted as number information, e.g. media server interpreted as form of virtual human service unit, Yang). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings of Yang into the teachings of Shan, to yield the result of the process behind generating and obtaining a virtual image in a video call session through a VoNR+ platform. Consider Claim 9. Yang teaches the video call control method of claim 8, wherein sending the virtual human image and number information of the first terminal device to the virtual human service unit comprises: sending the virtual human image (Pg. 10 lines 40-42, “The auxiliary media stream may also include one or more of point cloud data, spatial data (which may also be referred to as spatial pose data), user-view video, or a virtual model”, e.g. virtual mode interpreted as virtual human image, Yang), the number information of the first terminal device and video setting requirement information to the virtual human service unit (Pg. 5 lines 44-45, “the first terminal device sends the first data stream and the second data stream to the media server”, e.g. first data stream sent from first terminal device interpreted as number information, e.g. media server interpreted as form of virtual human service unit, Yang). Consider Claim 10. Shan teaches the video call control method of claim 9, wherein the video setting requirement information comprises one of: replacing a head image in the video with the virtual human image (Pg. 4 lines 6-8, “In the menu, the character image of the first user in the video call is replaced with a virtual image, which can be a full-body image replacement or only an avatar replacement”, e.g. replacing first user with virtual image interpreted as replacing head image in video, Shan) replacing a background in the video with the virtual human image (Pg. 5 lines 12-13, “…which includes the real background of the first user Augmented reality video of the avatar in the image”, e.g. augmented reality video of avatar includes the real background thus interpreted as replacing background with virtual human image, Shan); replacing a head image in the video with the virtual human image, and blurring a background in the video (Pg. 8 lines 22-24, “server generates an enhanced video image and sends it to the second user terminal, where the second user can see the augmented reality video during the video call with the first user on the second user terminal, which includes the real background of the…”, e.g. generating enhanced video image containing background is interpreted as replacing head image with virtual human image and performing a blurring processing on a background, Shan; or replacing a head image in the video with the virtual human image, and applying motion effects to a virtual human head image after the replacement (Pg 7 line 16-17, “Associate the facial expression information and body motion information with the virtual character skeleton to generate a virtual portrait model synchronized with the facial expression and body motion of the first user”, e.g. generating virtual portrait model associated with facial expression/body motion interpreted as replacing head image with virtual human image and applying motion effects to virtual human head, Shan). Consider Claim 11. Yang teaches the video call control method of claim 7, wherein the VoNR+ platform sending the virtual human video data to the second terminal device via the session access unit comprises: in response to the second terminal device being a terminal device which supports a data channel, sending, by the VoNR+ platform (Pg. 8 line 31, “and can also be used for a voice over new radio (Voice over New Radio, VoNR) video call.”, Yang), the virtual human video data to the second terminal device through a data transmission channel via the session access unit (See Pg. 2, SUMMARY OF THE INVENTION lines 18-19 “A video image, superimposing the user-drawn pattern on the user's video image to obtain a third video stream, and sending the third video stream to the second terminal device”, e.g. user’s video image sent to second terminal device interpreted as sending virtual human video data to the second terminal device, e.g. second terminal device contains data transmission channel whereas first and third terminal devices do not Yang); or, in response to the second terminal device being a terminal device which does not support data channel (Pg. 11 lines 19-20, “The terminal devices involved in the embodiments of this application, such as the first terminal device and the third terminal device, all support video calls, for example, support VoLTE”, e.g. second terminal device does not support video calls compared to first and third terminal device and video calls involves a data channel, Yang), performing, by the VoNR+ platform, a format conversion processing on the virtual human video data via the session access unit and sending the virtual human video data subjected to the format conversion processing to the second terminal device through a video transmission channel Pg. 27 lines 2-6, “As an example, the AME1 may encapsulate the superimposed video data into data packets to obtain the superimposed video stream, and the data packets obtained by encapsulating the superimposed video data may be in RTP or RTCP format”, e.g. superimposed video data interpreted as virtual human video data, e.g. encapsulating superimposed video data in RTP or RTCP format interpreted as performing format conversion processing on the data, e.g. AME1 interpreted as a session access unit, e.g. superimposed video stream interpreted as a video transmission channel, Yang). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings of Yang into the teachings of Shan, to yield the result of the process behind generating and obtaining a virtual image in a video call session through a VoNR+ platform. Consider Claim 12. Yang further teaches a communication device, comprising: a memory, a processor (Pg. 11 line 29, “The terminal device may include a processor 110, an external memory interface 120, an internal memory 121,”, Yang), and a computer program stored on the memory and executable by the processor (Pg. 14 liens 3-4, “Processor 110 may include one or more GPUs that execute program instructions to generate or alter display information.”, Yang), wherein the computer program, when executed by the processor, causes the processor to implement the video call control method of claim 1 (Pg. 14 lines 30-31, “The program storage area may store an operating system, an application program (such as a camera application) required for at least one function, and the like…“, e.g. application program interpreted as implementing the video call control method, Yang). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings of Yang into the teachings of Shan, to yield the result of the process behind generating and obtaining a virtual image in a video call session through a VoNR+ platform. Consider Claim 13. Yang teaches a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to cause a computer to implement the video call control method of claim 1 (Pg. 7 lines 15-20, “In an eighth aspect, the present application further provides a computer-readable storage medium, where instructions are stored in the computer-readable storage medium, when the computer-readable storage medium runs on a computer”, e.g. instructions from the computer readable storage medium lead to the implementation of video call control method, Yang). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings of Yang into the teachings of Shan, to yield the result of the process behind generating and obtaining a virtual image in a video call session through a VoNR+ platform. Consider Claim 14. Shan teaches acquiring video setting requirement information (Pg. 2 line 23, “where the virtual demand information is used to describe Image modification requirement”, e.g. image modification requirement interpreted as a video setting requirement information, Shan) of the first terminal device (Pg. 3 line 42-43, “after the first user connects to the second user through a user terminal including mobile phone applications, computer software, etc.”, e.g. first user connection via user terminal interpreted as first terminal device (connection requires a device), Shan) via the virtual human service unit (Pg. 5 line 38, “The video processing unit is configured to generate an enhanced video image”, e.g. video processing unit interpreted as virtual human service unit, Shan); and performing the replacement processing on the target object in the video according to the video setting requirement information and the virtual human image to obtain the virtual human video data (Pg. 2 line 23-25, “the model generation unit is configured to generate a virtual portrait model based on the personal image of the first user and the virtual requirement information; the video processing unit is configured to generate an enhanced video image, wherein The enhanced video image is obtained”, e.g. generating virtual portrait based on personal image interpreted as performing the replacement processing on target object, e.g. virtual requirement information interpreted as according to the video setting requirement, e.g. enhanced video image being obtained interpreted as obtaining virtual human video data, Shan). Consider Claim 15. Yang teaches acquiring, by the call capability module (Pg. 10 lines 28-29, “a proxy-CSCF (Proxy CSCF, P-CSCF), an interrogating-CSCF (Interrogating-CSCF, I-CSCF), or multiple “, e.g. I-CSCF reads on call capability module, Yang), the video call connection request (Pg. 24 line 23-24, “during the process of establishing a call connection with the second terminal device”, Yang) of the first terminal device via the call session control function unit (Pg. 10 lines 33-34, “…and can provide access control, quality of service control…”, e.g. access control interpreted as call session control function unit, Yang), and sending, by the call capability module, the video call connection request to the virtual human service unit via the service module; sending, by the service module, the virtual human image corresponding to the first terminal device acquired from the virtual human service unit to the media module (Pg. 10 lines 1-5, “the application server receives information from the first terminal device (such as call interface operation instructions). ), the application server sends the received information to the AR media server; thus, the AR media server processes the upstream media stream according to the information from the first terminal device”, e.g. AR media server interpreted as media module, Yang); and performing, by the media module, the replacement processing on the target object in the video with the virtual human image to obtain the virtual human video data (Pg. 8 lines 53-56, “The user video may include a video collected by a camera (a front-facing camera or a rear-facing camera) of a terminal device adopted by the user, or a file opened by the user, or an avatar of the user, and the like. For example, AR media enabler has strong image processing functions and data calculation functions, and can use AR technology to perform logical operations, screen rendering, virtual scene synthesis and other operations on the received media streams.”, e.g. screen rendering interpreted as replacement processing, e.g. avatar of user interpreted as virtual human image/virtual human video data, Yang), and sending, by the media module, the virtual human video data to the second terminal device via the session access unit (See Pg. 2, SUMMARY OF THE INVENTION lines 18-19 “A video image, superimposing the user-drawn pattern on the user's video image to obtain a third video stream, and sending the third video stream to the second terminal device”, e.g. user’s video image sent to second terminal device interpreted as sending virtual human video data to the second terminal device, Yang). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings of Yang into the teachings of Shan, to yield the result of the process behind generating and obtaining a virtual image in a video call session through a VoNR+ platform. Consider Claim 16. Yang teaches in response to the second terminal device being a terminal device which supports a data channel (“the terminal device may establish different auxiliary transmission channels according to the type of the auxiliary media stream to be transmitted”, e.g. terminal device interpreted as second terminal device being a terminal device, e.g. aux transmission channel interpreted as a data channel, Yang), sending the virtual human video data (Pg. 8 lines 53-56, “The user video may include a video collected by a camera (a front-facing camera or a rear-facing camera) of a terminal device adopted by the user, or a file opened by the user, or an avatar of the user, and the like.”, e.g. avatar of user interpreted as virtual human image/virtual human video data, Yang) to the session access unit (Pg. 10 lines 33-34, “…and can provide access control, quality of service control…”, e.g. access control interpreted as call session control function unit, Yang), so as to send the virtual human video data to the second terminal device through a data transmission channel (“the terminal device may establish different auxiliary transmission channels according to the type of the auxiliary media stream to be transmitted”, e.g. terminal device interpreted as second terminal device being a terminal device, e.g. aux transmission channel interpreted as a data channel, Yang) via the session access unit (Pg. 10 lines 33-34, “…and can provide access control, quality of service control…”, e.g. access control interpreted as call session control function unit, Yang); or, in response to the second terminal device being a terminal device which does not support data channel (Pg. 11 lines 19-20, “The terminal devices involved in the embodiments of this application, such as the first terminal device and the third terminal device, all support video calls, for example, support VoLTE”, e.g. second terminal device does not support video calls compared to first and third terminal device and video calls involves a data channel, Yang), sending the virtual human video data to the session access unit to perform a format conversion processing on the virtual human video data (Pg. 27 lines 2-6, “As an example, the AME1 may encapsulate the superimposed video data into data packets to obtain the superimposed video stream, and the data packets obtained by encapsulating the superimposed video data may be in RTP or RTCP format”, e.g. superimposed video data interpreted as virtual human video data, e.g. encapsulating superimposed vide data in RTP or RTCP format interpreted as performing format conversion processing on the data, e.g. AME1 interpreted as a session access unit, Yang) via the session access unit, and sending the virtual human video data subjected to the format conversion processing to the second terminal device through a video transmission channel (Pg. 27 lines 14-15 “910, The first terminal device displays the user-drawn pattern to be displayed. 911,The second terminal device displays the user-drawn pattern to be displayed and plays user audio data included in the audio data packet.”, e.g. displaying user-drawn pattern and user audio involves video transmission channel, Yang). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings of Yang into the teachings of Shan, to yield the result of the process behind generating and obtaining a virtual image in a video call session through a VoNR+ platform. Consider Claim 17. Yang teaches a memory, a processor (Pg. 11 line 29, “The terminal device may include a processor 110, an external memory interface 120, an internal memory 121,”, Yang), and a computer program stored on the memory and executable by the processor (Pg. 14 liens 3-4, “Processor 110 may include one or more GPUs that execute program instructions to generate or alter display information.”, Yang), wherein the computer program, when executed by the processor, causes the processor to implement the video call control method of claim 7 (Pg. 14 lines 30-31, “The program storage area may store an operating system, an application program (such as a camera application) required for at least one function, and the like…“, e.g. application program interpreted as implementing the video call control method, Yang). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings of Yang into the teachings of Shan, to yield the result of the process behind generating and obtaining a virtual image in a video call session through a VoNR+ platform. Consider Claim 18. Yang teaches sending the virtual human image (Pg. 10 lines 40-42, “The auxiliary media stream may also include one or more of point cloud data, spatial data (which may also be referred to as spatial pose data), user-view video, or a virtual model”, e.g. virtual mode interpreted as virtual human image, Yang) and number information of the first terminal device to the virtual human service unit (Pg. 5 lines 44-45, “the first terminal device sends the first data stream and the second data stream to the media server”, e.g. first data stream sent from first terminal device interpreted as number information, e.g. media server interpreted as form of virtual human service unit, Yang). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings of Yang into the teachings of Shan, to yield the result of the process behind generating and obtaining a virtual image in a video call session through a VoNR+ platform. Consider Claim 19. Yang teaches sending the virtual human image (Pg. 10 lines 40-42, “The auxiliary media stream may also include one or more of point cloud data, spatial data (which may also be referred to as spatial pose data), user-view video, or a virtual model”, e.g. virtual mode interpreted as virtual human image, Yang), the number information of the first terminal device and video setting requirement information to the virtual human service unit (Pg. 5 lines 44-45, “the first terminal device sends the first data stream and the second data stream to the media server”, e.g. first data stream sent from first terminal device interpreted as number information, e.g. media server interpreted as form of virtual human service unit, Yang). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings of Yang into the teachings of Shan, to yield the result of the process behind generating and obtaining a virtual image in a video call session through a VoNR+ platform. Consider Claim 20. Yang teaches a non-transitory computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to cause a computer to implement the video call control method of claim 7 (Pg. 7 lines 15-20, “In an eighth aspect, the present application further provides a computer-readable storage medium, where instructions are stored in the computer-readable storage medium, when the computer-readable storage medium runs on a computer”, e.g. computer-readable storage medium interpreted as a non-transitory computer-readable storage medium, e.g. instructions from the computer readable storage medium lead to the implementation of video call control method, Yang). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings of Yang into the teachings of Shan, to yield the result of the process behind generating and obtaining a virtual image in a video call session through a VoNR+ platform. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to VIET NGUYEN whose telephone number is (571)270-0174. The examiner can normally be reached 8am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached at (571) 272-7503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /VIET NGUYEN/Examiner, Art Unit 2691 /DUC NGUYEN/Supervisory Patent Examiner, Art Unit 2691
Read full office action

Prosecution Timeline

Jan 07, 2025
Application Filed
Sep 16, 2026
Non-Final Rejection mailed — §103 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month