Prosecution Insights
Last updated: October 01, 2026
Application No. 18/644,697

MULTIMODAL DEEPFAKE DETECTION VIA LIP-AUDIO CROSS-ATTENTION AND FACIAL SELF-ATTENTION

Final Rejection §103
Filed
Apr 24, 2024
Priority
Jun 27, 2023 — provisional 63/510,416
Examiner
DULANEY, KATHLEEN YUAN
Art Unit
2666
Tech Center
2600 — Communications
Assignee
Purdue Research Foundation
OA Round
2 (Final)
77%
Grant Probability
Favorable
3-4
OA Rounds
8m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
515 granted / 669 resolved
+15.0% vs TC avg
Strong +24% interview lift
Without
With
+24.3%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
25 currently pending
Career history
706
Total Applications
across all art units

Statute-Specific Performance

§101
9.8%
-30.2% vs TC avg
§103
40.1%
+0.1% vs TC avg
§102
17.2%
-22.8% vs TC avg
§112
30.1%
-9.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 669 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION The response received on 8/5/2026 has been placed in the file and was considered by the examiner. An action on the merit follows. Response to Amendment The amendments filed on 2026 August 5 have been fully considered. Response to these amendments is provided below. Summary of Amendment/ Arguments and Examiner’s Response: The applicant has amended the claims and has argued that prior art does not teach the newly claimed limitations. All arguments are moot in view of new grounds of rejection, below. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1 and 19 are rejected under 35 U.S.C. 103(a) as being unpatentable over U.S. Patent Application Publication NO. 20220121868 (Chen et al) in view of U.S. Patent Application Publication NO. 20230154040 (Guo et al) and U.S. Patent Application Publication No. 20220101087 (Lee et al). Regarding claim 1, Chen et al discloses a method (fig. 6, 8) for detecting whether a video has been manipulated, the method comprising: receiving, with a processor (fig. 2, item 202), a video (fig. 2, item 206) including a plurality of frames (fig. 2, item 210) and audio (fig. 2, item 208) of a person speaking (page 1, paragraph 10); determining, with the processor (page 10, paragraph 101), a first embedding (fig. 6, item 622) based on the plurality of frames of the video (fig. 6, item 606) using a first neural network (fig. 6, item 612), the first neural network being configured to detect artifacts in a facial region of the plurality of frames of the video by detecting if the frames match artifact spoofprint (page 5, paragraph 50-51); determining, with the processor (page 10, paragraph 101), a second embedding (fig. 6, item 618) based on the audio of the video (fig. 6, item 604) and the plurality of frames of the video (fig. 6, item 606) using a second neural network (fig. 6, item 610), the second neural network being configured to identify discrepancies between (i) lip movements of the person in the plurality of frames of the video and (ii) words spoken in the audio of the video (page 5, paragraph 54); and determining, with the processor (page 10, paragraph 101), whether the video has been manipulated based on both the first embedding and the second embedding (fig. 6, item 624, 626, page 6, paragraph 61). Chen et al further discloses that the audio is of the video (fig. 2), and that image data input by the first neural network is the plurality of frames of video (fig. 6, item 606). Chen et al does not disclose expressly the first neural network that process a facial region includes at least one self-attention layer that determines first attention weights based on values derived from the image data input, and that the second neural network that processes audio and video incorporates at least one cross-attention layer that determines second attention weights based on (i) values derived from the plurality of frames of the video and (ii) values derived from the audio. Guo et al discloses the first neural network (page 2, paragraph 19) that process a facial region (page 1, paragraph 2) includes at least one self-attention layer that determines first attention weights (page 2, paragraph 14) based on values derived from the image data (page 11, paragraph 130). Chen et al and Guo et al are combinable because they are from the same field of endeavor, i.e. neural network processing of facial images. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to use a self-attention layer. The suggestion/motivation for doing so would have been to provide a more robust method by using known methods for single input neural networks. Chen et al (as modified by Guo et al) does not disclose expressly that the second neural network that processes audio and video incorporates at least one cross-attention layer that determines second attention weights based on (i) values derived from the plurality of frames of the video and (ii) values derived from the audio. Lee et al discloses that the second neural network (page 1, paragraph 5) that processes audio and video (page 6, paragraph 56) incorporates at least one cross-attention layer (fig. 4a, item 434) that determines second attention weights (page 6, paragraph 56) based on (i) values derived from the plurality of frames of the video (page 6, paragraph 56) and (ii) values derived from the audio (page 6, paragraph 56). Chen et al (as modified by Guo et al) and Lee et al are combinable because they are from the same field of endeavor, i.e. neural network processing of audio and visual data. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to use a cross attention layer. The suggestion/motivation for doing so would have been to provide a more robust method by providing a means to process multiple inputs into the neural network. Therefore, it would have been obvious to combine the method of Chen et al with the self-attention layer of a neural network of Guo et al and the cross attention layer for audio and video inputs of Lee et al to obtain the invention as specified in claim 1. Regarding claim 19, Chen et al discloses the determining whether the video has been manipulated further comprising: determining a joint embedding by concatenating the first embedding and the second embedding (fig. 6, item 624, page 11, paragraph 106); and determining whether the video has been manipulated based on the joint embedding (fig. 6, item 624, page11, paragraph 107) using a linear neural network layer, the last layer of machine learning architecture fig. 6, item 607 of fig. 6, item 624 to fig. 6, item 626). Claims 2 and 4 are rejected under 35 U.S.C. 103(a) as being unpatentable over Chen et al in view of Guo et al and Lee et al, as applied to claim 1 above, and further in view of U.S. Patent No. 11373449 (Genner) Regarding claim 2, Chen et al (as modified by Guo et al and Lee et al) discloses all of the claimed elements as set forth above and incorporated herein by reference. Chen et al discloses the determining the first embedding further comprising processing the facial video/ images (fig. 6, item 512). Chen et al (as modified by Guo et al and Lee et al) does not disclose expressly processing the facial images includes generating a first plurality of cropped frames, each cropped frame of the first plurality of cropped frames being cropped by a bounding box around a face of the person in a respective frame of the plurality of frames. Genner discloses processing the facial video includes generating a first plurality of cropped frames (col. 10, lines 60-61), each cropped frame of the first plurality of cropped frames being cropped by a bounding box around a face of the person in a respective frame of the plurality of frames, a square (col 10, lines 61-62). Chen et al (as modified by Guo et al and Lee et al) and Genner are combinable because they are from the same field of endeavor, i.e. processing facial images. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to crop the facial area. The suggestion/motivation for doing so would have been to provide a more robust method by focusing on relevant data. Therefore, it would have been obvious to combine the method of Chen et al (as modified by Guo et al and Lee et al) with the cropping of Genner to obtain the invention as specified in claim 2. Regarding claim 4, Genner et al discloses generating the first plurality of cropped frames further comprising: resizing each of the first plurality of cropped frames to a first predetermined resolution (col 10, lines 61-62). Claim 3 is rejected under 35 U.S.C. 103(a) as being unpatentable over Chen et al in view of Guo et al, Lee et al, and Genner, as applied to claim 2 above, and further in view of U.S. Patent NO. 12266212 (Gupta). Regarding claim 3, Chen et al (as modified by Guo et al, Lee et al and Genner) discloses all of the claimed elements as set forth above and incorporated herein by reference. Chen et al (as modified by Guo et al, Lee et al and Genner) does not disclose expressly generating the first plurality of cropped frames further comprising: identifying a subset of the plurality of frames of the video; and generating the first plurality of cropped frames by cropping the subset of the plurality of frames of the video. Gupta et al discloses generating the first plurality of cropped frames further comprising: identifying a subset of the plurality of frames of the video (col. 8, lines 15-20); and generating the first plurality of cropped frames by cropping the subset of the plurality of frames of the video (col. 8, lines 40-41). Chen et al (as modified by Guo et al, Lee et al and Genner) and Gupta et al are combinable because they are from the same field of endeavor, i.e. processing facial images. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to use a subset of frames. The suggestion/motivation for doing so would have been to provide a faster method by processing only essential data. Therefore, it would have been obvious to combine the method of Chen et al (as modified by Guo et al, Lee et al and Genner) with a subset of frames of Gupta et al to obtain the invention as specified in claim 3. Claims 5, 7 and 8 are rejected under 35 U.S.C. 103(a) as being unpatentable over Chen et al in view of Guo et al, Lee et al and Genner, as applied to claim 2 above, and further in view of U.S. Patent Application Publication NO. 20230401824 (Khan et al). Regarding claim 5, Chen et al (as modified by Guo et al, Lee et al and Genner) discloses all of the claimed elements as set forth above and incorporated herein by reference. Genner discloses cropping the face from the image (col. 10, lines 60-61), and Chen et al discloses the first neural network is a network for facial features in video (fig. 6, item 612). Chen et al (as modified by Guo et al, Lee et al and Genner) does not disclose expressly determining the first embedding further comprising: determining a plurality of patches by extracting features from each of the first plurality of facial frames using a feature extractor of the facial network determining the first embedding further comprising: determining a plurality of patches. Khan et al discloses determining the first embedding further comprising: determining a plurality of patches (Fig. 8, item 812, 816) by extracting features from each of the first plurality of facial frames (fig. 6, item 602), the features that are extracted in fig. 8, item 812, or in fig. 8, item 814) using a feature extractor of the facial network, fig. 8 being part of the facial network. Chen et al (as modified by Guo et al, Lee et al and Genner) & Khan et al are combinable because they are from the same field of endeavor, i.e. detecting deepfakes. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to extract patches. The suggestion/motivation for doing so would have been to provide a more robust system by extracting features from significant areas of the image. Therefore, it would have been obvious to combine the method of Chen et al (as modified by Guo et al, Lee et al and Genner) with patch extraction of Khan et al to obtain the invention as specified in claim 5. Regarding claim 7, Khan et al discloses the determining the first embedding further comprising: determining the first embedding based on the plurality of patches using a Transformer encoder of the first neural network (fig. 8, item 808) having a multi-headed self-attention mechanism (fig. 8, item 826), which includes at least one self-attention layer (page 7, paragraph 59). Regarding claim 8, Khan et al discloses the determining the first embedding further comprising: determining a plurality of patch embeddings based on the plurality of patches using a linear layer (fig. 8, item 814); and determining a first plurality of position-encoded embeddings by embedding position information into the plurality of patch embeddings (fig. 8, item 806), wherein the first embedding is determined based on the first plurality of position-encoded embeddings (fig. 8, output of transform encoder 800 is based on item 816). Claims 11-14 are rejected under 35 U.S.C. 103(a) as being unpatentable over Chen et al in view of Guo et al and Lee et al, as applied to claim 1 above, and further in view of U.S. Patent Application Publication No. 20220318349 (Wasnik et al). Regarding claim 11, Chen et al (as modified by Guo et al and Lee et al) discloses all of the claimed elements as set forth above and incorporated herein by reference. Chen et al further discloses determining the second embedding further comprises analyzing a bounding box around lips of the person in a respective frame of the plurality of frames (page 5, paragraph 56). Chen et al (as modified by Guo et al and Lee et al) does not disclose expressly generating a second plurality of cropped frames, each cropped frame of the second plurality of cropped frames being cropped by a bounding box around lips of the person. Wasnik et al discloses generating a second plurality of cropped frames, each cropped frame of the second plurality of cropped frames being cropped by a bounding box around lips of the person, around a mouth region (page 7, paragraph 85). Chen et al (as modified by Guo et al and Lee et al) and Wasnik et al are combinable because they are from the same field of endeavor, i.e. analyzing speech in videos. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to crop the mouth area. The suggestion/motivation for doing so would have been to provide a faster system by lowering the data to be processed. Therefore, it would have been obvious to combine the method of Chen et al (as modified by Guo et al and Lee et al) with the cropping of Wasnik et al to obtain the invention as specified in claim 11. Regarding claim 12, Wasnik et al discloses the generating the second plurality of cropped frames further comprising: resizing each of the second plurality of cropped frames to a second predetermined resolution, 60x100 pixels (page 7, paragraph 85). Regarding claim 13, Wasnik et al discloses generating the second plurality of cropped frames further comprising: converting the second plurality of cropped frames to greyscale (page 7, paragraph 85). Regarding claim 14, Wasnik et al discloses converting the audio of the video into mono audio, i.e. a single channel stream (page 7, paragraph 86). Allowable Subject Matter Claim 20 is allowed. Claim 20 incorporated previously indicated allowable subject matter Claims 6, 9, 10, 15-18 are objected to as being dependent upon a rejected base claim but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Claim 15 contains allowable subject matter regarding determining the claimed second embedding based on the claimed second plurality of cropped frames cropped by a bounding box around lips of the person, and the audio of the video using a Transformer encoder of the second neural network having a cross attention mechanism, the cross-attention mechanism including the claimed at least one cross-attention layer that determines second attention weights based on the claimed values. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Kathleen Yuan Dulaney whose telephone number is (571)272-2902. The examiner can normally be reached M-F: 9AM-5PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at 5712703717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KATHLEEN Y DULANEY/ Primary Examiner, Art Unit 2666 9/8/2026
Read full office action

Prosecution Timeline

Apr 24, 2024
Application Filed
May 07, 2026
Non-Final Rejection mailed — §103
Aug 05, 2026
Response Filed
Sep 16, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12731244
INSPECTION SYSTEM AND METHOD FOR CONTROLLING INSPECTION SYSTEM
3y 1m to grant Granted Sep 08, 2026
Patent 12731247
DEBRIS DETERMINATION METHOD
2y 9m to grant Granted Sep 08, 2026
Patent 12725272
NON-TRANSITORY COMPUTER READABLE MEDIUM AND COMMUNICATION METHOD
3y 3m to grant Granted Sep 01, 2026
Patent 12725407
LIGHT BACKDOOR ATTACK METHOD IN PHYSICAL WORLD
2y 8m to grant Granted Sep 01, 2026
Patent 12688579
REGISTRATION OF 3D and 2D IMAGES FOR SURGICAL NAVIGATION AND ROBOTIC GUIDANCE WITHOUT USING RADIOPAQUE FIDUCIALS IN THE IMAGES
3y 3m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
77%
Grant Probability
99%
With Interview (+24.3%)
3y 1m (~8m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 669 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month