Prosecution Insights
Last updated: September 17, 2026
Application No. 18/747,370

EMOTIONAL ENGAGEMENT DETECTION METHOD BASED ON POSITIVE EMOTIONAL PERCEPTION

Final Rejection §103
Filed
Jun 18, 2024
Priority
Jul 04, 2023 — CN 202310811417.5
Examiner
LIN, JESSICA YIFANG
Art Unit
2668
Tech Center
2600 — Communications
Assignee
Jiangxi Normal University
OA Round
2 (Final)
82%
Grant Probability
Favorable
3-4
OA Rounds
2m
Est. Remaining
78%
With Interview

Examiner Intelligence

Grants 82% — above average
82%
Career Allowance Rate
9 granted / 11 resolved
+19.8% vs TC avg
Minimal -3% lift
Without
With
+-3.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 5m
Avg Prosecution
58 currently pending
Career history
63
Total Applications
across all art units

Statute-Specific Performance

§101
3.1%
-36.9% vs TC avg
§103
63.8%
+23.8% vs TC avg
§102
29.6%
-10.4% vs TC avg
§112
3.1%
-36.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 11 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Applicant’s amendment filed 6/30/2026 have been entered and made of record. Applicant has amended claim 1. Claims 2-5 remain unchanged. Claim 6 is newly added. No new matter is entered. Currently, claims 1-6 are being addressed. Response to Arguments Applicant argues that the prior arts of record does not relate to classroom teaching scenarios, does not involve any learning system, and does not address emotional engagement detection. Furthermore, applicant argues the cited prior art references do not disclose any fusion architecture at all—whether decision fusion, multimodal fusion, or a combination thereof. The applicant argues that the prior art in combination fails to disclose, teach or suggest the dual path fusion process as recited in currently amended claim 1 of the instant application. In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). Applicant's arguments filed 6/30/2026 have been fully considered but they are not persuasive. Examiner has performed an updated search and new prior art was found to teach the feature added to the now amended claim 1, “to obtain a thinking activity, wherein the access records comprise exploration, access, click, and comment information of the classroom student on teaching resources in the learning system.” Furthermore, claim 6 is substantially the same as claim 1 except for “wherein the deep learning network corresponding to the smile intensity estimation model comprises: a VGGNet and a ResNet; for the VGGNet, a pre-trained weight is loaded, the pre-trained weight is taken as a starting point for learning to compensate for a shortage of smile datasets, and to capture smile detailed information; and for the ResNet, a pre-trained weight is not loaded, an image feature is learned from scratch, and the ResNet is configured to focus on training information of the smile intensity; constructing a multi-task head pose estimation model by using a FSANet”. Upon further consideration, a new ground(s) of rejection is made in view of Smith (United States Patent Application Publication US 2021/0049488 A1) and Yu (Chinese Patent CN 108805089 B). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claim(s) 1-4 are rejected under 35 U.S.C. 103 as being unpatentable over Zhenzhen Luo et al., "Algorithm Modeling of Smile IntensityEstimation Using a Spatial Attention Convolutional Neural Network", "Journal of Physics:Conference Series", Vol. 1746 No. 1, 1-12, 20210131 in view of Liu et. al. (Chinese patent CN-116244474-A), Chen et. al. (Chinese patent CN-108399376-A), Wen et. al. (Chinese patent CN-112686195-A), Hazur et. al. (United States Patent Application Publication US-20180018540-A1) and Smith (United States Patent Application Publication US 2021/0049488 A1). Regarding claim 1, Zhenzhen et. al. discloses an emotional engagement detection method based on positive emotional perception, comprising: obtaining a head pose direction, a smile intensity; constructing a smile intensity estimation model by introducing an attention mechanism in a fine-grained image recognition to a deep learning network, and inputting the smile intensity of the classroom student into the smile intensity estimation model to obtain an emotional intensity (Zhenzhen et. al. section 1-4 of the body, VGGNet16 and VGGNet19 used to build a smiley face intensity estimation model, Table 1 shows model parameters equivalent to introducing an attention mechanism in a deep learning network into fine-grained image recognition); constructing a multi-task head pose estimation model by using a feature-and-spatial aligned network (FSANet) It is broadly interpreted that the images used can be faces of a classroom student). However, Zhenzhen et. al. fails to disclose access records of a learning system of a classroom student; inputting the head pose direction of the classroom student into the multi-task head pose estimation model to obtain a cognitive attention; screening the access records of the learning system of the classroom student to obtain a thinking activity, wherein the access records comprise exploration, access, click, and comment information of the classroom student on teaching resources in the learning system; constructing a positive emotional engagement recognition model based on decision fusion of the classroom student according to a positive emotional engagement model by using the cognitive attention, the emotional intensity and the thinking activity; obtaining classroom video images and terminal system data for teaching resources, inputting the head pose direction and the smile intensity of the classroom student in the classroom video images and the access records of the learning system in the terminal system data for teaching resources, which are considered as information of complementary different modalities, into a deep network, to perform high-level semantic learning on the different modalities and perform feature fusion on the different modalities, to thereby obtain homogeneous fusion expression; formulating, according to the head pose direction, the smile intensity and the access records of the learning system of the classroom student, different objective functions, and learning the homogeneous fusion expression through the different objective functions to thereby construct a positive emotional engagement recognition model based on multimodal fusion of the classroom student; constructing a positive emotional engagement recognition model of the classroom student by combining the positive emotional engagement recognition model based on decision fusion of the classroom student and the positive emotional engagement recognition model based on multimodal fusion of the classroom student; and recognizing the smile intensity, the head pose direction and the cognitive attention by using the positive emotional engagement recognition model of the classroom student to obtain emotional engagement based on positive emotional perception of the classroom student. Wen et. al. teaches inputting the head pose direction of the classroom student into the multi-task head pose estimation model to obtain a cognitive attention (Wen et. al., see [0125]-[0126] of the specification, the terminal uses the head pose model FSANet to extract features on the input face picture, and uses the extracted features to extract three Euler angles (pitch, yaw, and roll) that infer the current face); Hazur et. al. teaches screening the access records of the learning system of the classroom student to obtain a thinking activity (Hazur et. al., [0046]-[0047]); constructing a positive emotional engagement recognition model based on decision fusion of the classroom student according to a positive emotional engagement model by using the cognitive attention, the emotional intensity and the thinking activity (Hazur et. al., [0062]-[0065]); Liu et. al. teaches obtaining classroom video images and terminal system data for teaching resources, inputting the head pose direction and the smile intensity of the classroom student in the classroom video images and the access records of the learning system in the terminal system data for teaching resources, which are considered as information of complementary different modalities, into a deep network, to perform high-level semantic learning on the different modalities and perform feature fusion on the different modalities, to thereby obtain homogeneous fusion expression (Liu et. al., page 5, paragraph 7-17); Chen et. al. teaches formulating, according to the head pose direction, the smile intensity and the access records of the learning system of the classroom student, different objective functions, and learning the homogeneous fusion expression through the different objective functions to thereby construct a positive emotional engagement recognition model based on multimodal fusion of the classroom student (Chen et. al., page 9-10, multi-modal fusion analytic learning interest from face detection and analysis module, classroom interactions’ cloud platform module, learning interest analysis module, estimation of head pose and expression, and concrete implementation mode); constructing a positive emotional engagement recognition model of the classroom student by combining the positive emotional engagement recognition model based on decision fusion of the classroom student and the positive emotional engagement recognition model based on multimodal fusion of the classroom student (Chen et. al., page 9-10, multi-modal fusion analytic learning interest from face detection and analysis module, classroom interactions’ cloud platform module, learning interest analysis module, estimation of head pose and expression, and concrete implementation mode); and recognizing the smile intensity, the head pose direction and the cognitive attention by using the positive emotional engagement recognition model of the classroom student to obtain emotional engagement based on positive emotional perception of the classroom student (Chen et. al., page 9-10, multi-modal fusion analytic learning interest from face detection and analysis module, classroom interactions’ cloud platform module, learning interest analysis module, estimation of head pose and expression, and concrete implementation mode, specific implementation mode: page 8, the positive negative-morality of student is judged according to expression). Smith teaches wherein the access records comprise exploration, access, click, and comment information of the classroom student on teaching resources in the learning system (Smith Figure 1, 2, 7C, 8, 9, 10, 11, [0005]: the communication model is further configured to manage bi-directional communication with the respective student based on matching a current communication and context to a communication option in the trained model. [0023]: Various embodiments of a monitoring and intervention system address problems with managing student engagement, and further embodiments manage the need to identify and resolve underlying causes for problem behavior and other schooling issues (e.g., attendance, etc.). PNG media_image1.png 466 654 media_image1.png Greyscale The claimed invention requires all the components and distinguishing features disclosed as part of their comprehensive approach to construct a positive emotional engagement recognition model of classroom students from multi-modal learning fusion as well as decision fusion. The technical solution claimed by this claim is obtained on the basis of Zhenzhen et. al. in combination with Wen et. al., Hazur et. al., Liu et. al., Chen et. al., and Smith. The inventive concept is the multi-tasking head pose estimation model using FSANet networks, with each of the prior arts providing technical insights to the solution provided by the claimed invention. Thus, it would have been obvious to one skilled in the art prior to the effective filing date of the claimed invention to have combined the teachings of Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al. and Smith to arrive at the emotional engagement detection method that is disclosed in the original application. Regarding claim 2, the combination of Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al. and Smith discloses the emotional engagement detection method based on positive emotional perception as claimed in claim 1. Zhenzhen et. al. further discloses wherein the smile intensity estimation model is configured to: based on the attention mechanism in the fine-grained image recognition, suppress useless information learned by a convolutional layer of the deep learning network, and enhance learning of features in a key area by the deep learning network; and the key area is an area where facial muscles that produce a smile are located. (Zhenzhen et. al., see section 3.2 of the body) Regarding claim 3, the combination of Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al. and Smith discloses the emotional engagement detection method based on positive emotional perception as claimed in claim 2, and Zhenzhen et. al. further discloses wherein features of the area where the facial muscles that produce the smile are located comprise: a coordinate change of a mouth corner feature point and a coordinate change of an eye corner feature point during a smile movement; and wherein the deep learning network corresponding to the smile intensity estimation model comprises: a visual geometry group network (VGGNet) and a residual neural network (ResNet); for the VGGNet, a pre-trained weight is loaded, the pre-trained weight is taken as a starting point for learning to compensate for a shortage of smile datasets, and to capture smile detailed information; and for the ResNet, a pre-trained weight is not loaded, an image feature is learned from scratch, and the ResNet is configured to focus on training information of the smile intensity. (Zhenzhen et. al., section 3.2 of body, VGGNet equivalent to deep networks with pre-trained weights, taking feature weights as starting points for learning, making up the disadvantage of the small number of smiley face datasets, and capturing detailed information about the smiley face, model parameters shown in Table 1) Regarding claim 4, the combination of Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al. and Smith discloses the emotional engagement detection method based on positive emotional perception as claimed in claim 3, and Zhenzhen et. al. further discloses further comprising: introducing a focal loss function into the deep learning network corresponding to the smile intensity estimation model to improve a cross-entropy loss function. (Zhenzhen et. al., see section 3.2 of body, the loss function consists of a function of the cross-entropy function and the ranking part used to distinguish the smiley face region) Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Zhenzhen Luo et al., "Algorithm Modeling of Smile Intensity Estimation Using a Spatial Attention Convolutional Neural Network", "Journal of Physics: Conference Series", Vol. 1746 No. 1, 1-12, 20210131 in view of Liu et. al. (Chinese patent CN-116244474-A), Chen et. al. (Chinese patent CN-108399376-A), Wen et. al. (Chinese patent CN-112686195-A), Hazur et. al. (United States Patent Application Publication US-20180018540-A1) and Smith (United States Patent Application Publication US 2021/0049488 A1), as applied to claim 1 above, and further in view of Yang et. al. (Chinese Patent CN114549439 A). Regarding claim 5, the combination of Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al. and Smith discloses the emotional engagement detection method based on positive emotional perception as claimed in claim 1. However, Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al. and Smith fail to disclose wherein the inputting the head pose direction and the smile intensity of the classroom student in the classroom video images and the access records of the learning system in the terminal system data for teaching resources, which are considered as information of complementary different modalities, into a deep network, to perform high-level semantic learning on the different modalities and perform feature fusion on the different modalities, to thereby obtain homogeneous fusion expression, specifically comprises: performing feature learning through asymmetric dual stream branches, wherein the asymmetric dual stream branches comprise a first branch and a second branch; the first branch is configured to utilize a deep cascaded autoencoder to perform feature mapping learning based on the head pose direction and the smile intensity of the classroom student in the classroom video images and the access records of the learning system through the deep network to achieve the high-level semantic learning; and the second branch is configured to extract low-level detail information for each of information of the different modalities; and connecting a generated higher-level feature of each of information of the different modalities to a lower-level feature of each of information of the different modalities through a skip layer, and generating a shared sparse fusion feature through an attention-guided feature cross fusion module, to correlate features of the different modalities together in a nonlinear manner and focus the features of the different modalities, and thereby obtain the homogeneous fusion expression. Yang et. al. teaches wherein the inputting the head pose direction and the smile intensity of the classroom student in the classroom video images and the access records of the learning system in the terminal system data for teaching resources, which are considered as information of complementary different modalities, into a deep network, to perform high-level semantic learning on the different modalities and perform feature fusion on the different modalities, to thereby obtain homogeneous fusion expression, specifically comprises: performing feature learning through asymmetric dual stream branches, wherein the asymmetric dual stream branches comprise a first branch and a second branch; the first branch is configured to utilize a deep cascaded autoencoder to perform feature mapping learning based on the head pose direction and the smile intensity of the classroom student in the classroom video images and the access records of the learning system through the deep network to achieve the high-level semantic learning; and the second branch is configured to extract low-level detail information for each of information of the different modalities; and connecting a generated higher-level feature of each of information of the different modalities to a lower-level feature of each of information of the different modalities through a skip layer, and generating a shared sparse fusion feature through an attention-guided feature cross fusion module, to correlate features of the different modalities together in a nonlinear manner and focus the features of the different modalities, and thereby obtain the homogeneous fusion expression (Yang et. al., see specification paragraphs [0007]-[0036], discloses an RBG-D image semantic segmentation method based on multimodal feature fusion, an RGB-D image semantic segmentation method based on multimodal feature fusion). These neural network features are critical to the claimed invention because it allows images to be inputted and subsequently categorized based on extracted sematic features. Yang et. al. describes the use of an input attention-guided multi-modal cross-fusion segmentation network model (ACFNe), which follows the encoder-decoder structure, with the ResNet-101 and ResNet-50 networks as backbone networks. All feature learning through asymmetric dual-stream branching, the high-level features of each information modality generated are connected by hop layers with features of their lower layers, by the attention-guided feature cross-fusion model, shared sparsely fused features are generated, and features of multiple modalities are correlated together in a complex non-linear manner to obtain a homogenous fused expression, using the fused augmented feature representation. Thus, it would have been obvious to one skilled in the art prior to the effective filing date of the claimed invention to have combined the teachings of Yang et. al. with the combination of Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al. and Smith to arrive at the solution presented by the claimed invention. Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhenzhen Luo et al., "Algorithm Modeling of Smile IntensityEstimation Using a Spatial Attention Convolutional Neural Network", "Journal of Physics:Conference Series", Vol. 1746 No. 1, 1-12, 20210131 in view of Liu et. al. (Chinese patent CN-116244474-A), Chen et. al. (Chinese patent CN-108399376-A), Wen et. al. (Chinese patent CN-112686195-A), Hazur et. al. (United States Patent Application Publication US-20180018540-A1) and Smith (United States Patent Application Publication US 2021/0049488 A1) as applied to claim 1 above, and further in view of Yu et. al. (Chinese Patent CN 108805089 B). Regarding claim 6, the subject matter is substantially the same as claim 1 except for “wherein the deep learning network corresponding to the smile intensity estimation model comprises: a VGGNet and a ResNet; for the VGGNet, a pre-trained weight is loaded, the pre-trained weight is taken as a starting point for learning to compensate for a shortage of smile datasets, and to capture smile detailed information; and for the ResNet, a pre-trained weight is not loaded, an image feature is learned from scratch, and the ResNet is configured to focus on training information of the smile intensity; constructing a multi-task head pose estimation model by using a FSANet”. Thus, the rejection of claim 1 is incorporated herein. Yu et. al. discloses wherein the deep learning network corresponding to the smile intensity estimation model comprises: a VGGNet and a ResNet; for the VGGNet, a pre-trained weight is loaded, the pre-trained weight is taken as a starting point for learning to compensate for a shortage of smile datasets, and to capture smile detailed information; and for the ResNet, a pre-trained weight is not loaded, an image feature is learned from scratch, and the ResNet is configured to focus on training information of the smile intensity; constructing a multi-task head pose estimation model by using a FSANet (Yang et. al. page 5, paragraphs 2-4: a multi-modal-based emotion recognition method based on facial image expression with fusion of VGG16 and RESNET 50 neural network architecture). PNG media_image2.png 764 1290 media_image2.png Greyscale This fusion neural network architecture is important to the claimed invention because it enables information extraction based on different objective functions, such as thinking activity of the student, the emotional state and cognitive attention of the student, which the neural network can train further for a subsequent learning task or used in the predictive model. The inventive concept is the multi-tasking head pose estimation model using FSANet networks, with each of the prior arts providing technical insights to the solution provided by the claimed invention. Thus, it would have been obvious to one skilled in the art prior to the effective filing date of the claimed invention to have combined the teachings of Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al., Smith and Yu et. al. to arrive at the emotional engagement detection method that is disclosed in the original application. Conclusion Response to Amendment Examiner has carefully considered the amended claims and performed an updated search. New prior art was found to capture the features included in the amended claims. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JESSICA YIFANG LIN whose telephone number is (571)272-6435. The examiner can normally be reached M-F 7:00am-6:15pm, with optional day off. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached at 571-272-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JESSICA YIFANG LIN/Examiner, Art Unit 2668 August 6, 2026 /ANAND P BHATNAGAR/Primary Examiner, Art Unit 2668 August 6, 2026
Read full office action

Prosecution Timeline

Jun 18, 2024
Application Filed
Mar 30, 2026
Non-Final Rejection mailed — §103
Jun 30, 2026
Response Filed
Aug 10, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12738010
OBJECT RECOGNITION DEVICE AND OBJECT RECOGNITION METHOD
2y 5m to grant Granted Sep 15, 2026
Patent 12718396
AVIATION DOCUMENT TARGET AREA EXTRACTION SYSTEM AND METHOD
2y 6m to grant Granted Aug 25, 2026
Patent 12711738
Real-time Media Alteration Using Generative Techniques
2y 9m to grant Granted Aug 18, 2026
Patent 12678109
CONTROL METHOD AND CONTROL SYSTEM FOR IMAGE SCANNING, ELECTRONIC APPARATUS, AND STORAGE MEDIUM
2y 8m to grant Granted Jul 14, 2026
Patent 12597139
CONTROLLING AN ALERT SIGNAL FOR SPECTRAL COMPUTED TOMOGRAPHY IMAGING
2y 3m to grant Granted Apr 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
82%
Grant Probability
78%
With Interview (-3.3%)
2y 5m (~2m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 11 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month