Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Applicant’s amendment filed 6/30/2026 have been entered and made of record. Applicant has amended claim 1. Claims 2-5 remain unchanged. Claim 6 is newly added. No new matter is entered. Currently, claims 1-6 are being addressed.
Response to Arguments
Applicant argues that the prior arts of record does not relate to classroom teaching scenarios, does not involve any learning system, and does not address emotional engagement detection. Furthermore, applicant argues the cited prior art references do not disclose any fusion architecture at all—whether decision fusion, multimodal fusion, or a combination thereof. The applicant argues that the prior art in combination fails to disclose, teach or suggest the dual path fusion process as recited in currently amended claim 1 of the instant application.
In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986).
Applicant's arguments filed 6/30/2026 have been fully considered but they are not persuasive. Examiner has performed an updated search and new prior art was found to teach the feature added to the now amended claim 1, “to obtain a thinking activity, wherein the access records comprise exploration, access, click, and comment information of the classroom student on teaching resources in the learning system.” Furthermore, claim 6 is substantially the same as claim 1 except for “wherein the deep learning network corresponding to the smile intensity estimation model comprises: a VGGNet and a ResNet; for the VGGNet, a pre-trained weight is loaded, the pre-trained weight is taken as a starting point for learning to compensate for a shortage of smile datasets, and to capture smile detailed information; and for the ResNet, a pre-trained weight is not loaded, an image feature is learned from scratch, and the ResNet is configured to focus on training information of the smile intensity; constructing a multi-task head pose estimation model by using a FSANet”. Upon further consideration, a new ground(s) of rejection is made in view of Smith (United States Patent Application Publication US 2021/0049488 A1) and Yu (Chinese Patent CN 108805089 B).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-4 are rejected under 35 U.S.C. 103 as being unpatentable over Zhenzhen Luo et al., "Algorithm Modeling of Smile IntensityEstimation Using a Spatial Attention Convolutional Neural Network", "Journal of Physics:Conference Series", Vol. 1746 No. 1, 1-12, 20210131 in view of Liu et. al. (Chinese patent CN-116244474-A), Chen et. al. (Chinese patent CN-108399376-A), Wen et. al. (Chinese patent CN-112686195-A), Hazur et. al. (United States Patent Application Publication US-20180018540-A1) and Smith (United States Patent Application Publication US 2021/0049488 A1).
Regarding claim 1, Zhenzhen et. al. discloses an emotional engagement detection method based on positive emotional perception, comprising: obtaining a head pose direction, a smile intensity; constructing a smile intensity estimation model by introducing an attention mechanism in a fine-grained image recognition to a deep learning network, and inputting the smile intensity of the classroom student into the smile intensity estimation model to obtain an emotional intensity (Zhenzhen et. al. section 1-4 of the body, VGGNet16 and VGGNet19 used to build a smiley face intensity estimation model, Table 1 shows model parameters equivalent to introducing an attention mechanism in a deep learning network into fine-grained image recognition); constructing a multi-task head pose estimation model by using a feature-and-spatial aligned network (FSANet) It is broadly interpreted that the images used can be faces of a classroom student).
However, Zhenzhen et. al. fails to disclose access records of a learning system of a classroom student; inputting the head pose direction of the classroom student into the multi-task head pose estimation model to obtain a cognitive attention;
screening the access records of the learning system of the classroom student to obtain a thinking activity, wherein the access records comprise exploration, access, click, and comment information of the classroom student on teaching resources in the learning system;
constructing a positive emotional engagement recognition model based on decision fusion of the classroom student according to a positive emotional engagement model by using the cognitive attention, the emotional intensity and the thinking activity;
obtaining classroom video images and terminal system data for teaching resources, inputting the head pose direction and the smile intensity of the classroom student in the classroom video images and the access records of the learning system in the terminal system data for teaching resources, which are considered as information of complementary different modalities, into a deep network, to perform high-level semantic learning on the different modalities and perform feature fusion on the different modalities, to thereby obtain homogeneous fusion expression;
formulating, according to the head pose direction, the smile intensity and the access records of the learning system of the classroom student, different objective functions, and learning the homogeneous fusion expression through the different objective functions to thereby construct a positive emotional engagement recognition model based on multimodal fusion of the classroom student;
constructing a positive emotional engagement recognition model of the classroom student by combining the positive emotional engagement recognition model based on decision fusion of the classroom student and the positive emotional engagement recognition model based on multimodal fusion of the classroom student;
and recognizing the smile intensity, the head pose direction and the cognitive attention by using the positive emotional engagement recognition model of the classroom student to obtain emotional engagement based on positive emotional perception of the classroom student.
Wen et. al. teaches inputting the head pose direction of the classroom student into the multi-task head pose estimation model to obtain a cognitive attention (Wen et. al., see [0125]-[0126] of the specification, the terminal uses the head pose model FSANet to extract features on the input face picture, and uses the extracted features to extract three Euler angles (pitch, yaw, and roll) that infer the current face);
Hazur et. al. teaches screening the access records of the learning system of the classroom student to obtain a thinking activity (Hazur et. al., [0046]-[0047]);
constructing a positive emotional engagement recognition model based on decision fusion of the classroom student according to a positive emotional engagement model by using the cognitive attention, the emotional intensity and the thinking activity (Hazur et. al., [0062]-[0065]);
Liu et. al. teaches obtaining classroom video images and terminal system data for teaching resources, inputting the head pose direction and the smile intensity of the classroom student in the classroom video images and the access records of the learning system in the terminal system data for teaching resources, which are considered as information of complementary different modalities, into a deep network, to perform high-level semantic learning on the different modalities and perform feature fusion on the different modalities, to thereby obtain homogeneous fusion expression (Liu et. al., page 5, paragraph 7-17);
Chen et. al. teaches formulating, according to the head pose direction, the smile intensity and the access records of the learning system of the classroom student, different objective functions, and learning the homogeneous fusion expression through the different objective functions to thereby construct a positive emotional engagement recognition model based on multimodal fusion of the classroom student (Chen et. al., page 9-10, multi-modal fusion analytic learning interest from face detection and analysis module, classroom interactions’ cloud platform module, learning interest analysis module, estimation of head pose and expression, and concrete implementation mode);
constructing a positive emotional engagement recognition model of the classroom student by combining the positive emotional engagement recognition model based on decision fusion of the classroom student and the positive emotional engagement recognition model based on multimodal fusion of the classroom student (Chen et. al., page 9-10, multi-modal fusion analytic learning interest from face detection and analysis module, classroom interactions’ cloud platform module, learning interest analysis module, estimation of head pose and expression, and concrete implementation mode);
and recognizing the smile intensity, the head pose direction and the cognitive attention by using the positive emotional engagement recognition model of the classroom student to obtain emotional engagement based on positive emotional perception of the classroom student (Chen et. al., page 9-10, multi-modal fusion analytic learning interest from face detection and analysis module, classroom interactions’ cloud platform module, learning interest analysis module, estimation of head pose and expression, and concrete implementation mode, specific implementation mode: page 8, the positive negative-morality of student is judged according to expression).
Smith teaches wherein the access records comprise exploration, access, click, and comment information of the classroom student on teaching resources in the learning system (Smith Figure 1, 2, 7C, 8, 9, 10, 11, [0005]: the communication model is further configured to manage bi-directional communication with the respective student based on matching a current communication and context to a communication option in the trained model. [0023]: Various embodiments of a monitoring and intervention system address problems with managing student engagement, and further embodiments manage the need to identify and resolve underlying causes for problem behavior and other schooling issues (e.g., attendance, etc.).
PNG
media_image1.png
466
654
media_image1.png
Greyscale
The claimed invention requires all the components and distinguishing features disclosed as part of their comprehensive approach to construct a positive emotional engagement recognition model of classroom students from multi-modal learning fusion as well as decision fusion. The technical solution claimed by this claim is obtained on the basis of Zhenzhen et. al. in combination with Wen et. al., Hazur et. al., Liu et. al., Chen et. al., and Smith. The inventive concept is the multi-tasking head pose estimation model using FSANet networks, with each of the prior arts providing technical insights to the solution provided by the claimed invention. Thus, it would have been obvious to one skilled in the art prior to the effective filing date of the claimed invention to have combined the teachings of Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al. and Smith to arrive at the emotional engagement detection method that is disclosed in the original application.
Regarding claim 2, the combination of Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al. and Smith discloses the emotional engagement detection method based on positive emotional perception as claimed in claim 1. Zhenzhen et. al. further discloses wherein the smile intensity estimation model is configured to: based on the attention mechanism in the fine-grained image recognition, suppress useless information learned by a convolutional layer of the deep learning network, and enhance learning of features in a key area by the deep learning network; and the key area is an area where facial muscles that produce a smile are located. (Zhenzhen et. al., see section 3.2 of the body)
Regarding claim 3, the combination of Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al. and Smith discloses the emotional engagement detection method based on positive emotional perception as claimed in claim 2, and Zhenzhen et. al. further discloses wherein features of the area where the facial muscles that produce the smile are located comprise: a coordinate change of a mouth corner feature point and a coordinate change of an eye corner feature point during a smile movement; and wherein the deep learning network corresponding to the smile intensity estimation model comprises: a visual geometry group network (VGGNet) and a residual neural network (ResNet); for the VGGNet, a pre-trained weight is loaded, the pre-trained weight is taken as a starting point for learning to compensate for a shortage of smile datasets, and to capture smile detailed information; and for the ResNet, a pre-trained weight is not loaded, an image feature is learned from scratch, and the ResNet is configured to focus on training information of the smile intensity. (Zhenzhen et. al., section 3.2 of body, VGGNet equivalent to deep networks with pre-trained weights, taking feature weights as starting points for learning, making up the disadvantage of the small number of smiley face datasets, and capturing detailed information about the smiley face, model parameters shown in Table 1)
Regarding claim 4, the combination of Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al. and Smith discloses the emotional engagement detection method based on positive emotional perception as claimed in claim 3, and Zhenzhen et. al. further discloses further comprising: introducing a focal loss function into the deep learning network corresponding to the smile intensity estimation model to improve a cross-entropy loss function. (Zhenzhen et. al., see section 3.2 of body, the loss function consists of a function of the cross-entropy function and the ranking part used to distinguish the smiley face region)
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Zhenzhen Luo et al., "Algorithm Modeling of Smile Intensity Estimation Using a Spatial Attention Convolutional Neural Network", "Journal of Physics: Conference Series", Vol. 1746 No. 1, 1-12, 20210131 in view of Liu et. al. (Chinese patent CN-116244474-A), Chen et. al. (Chinese patent CN-108399376-A), Wen et. al. (Chinese patent CN-112686195-A), Hazur et. al. (United States Patent Application Publication US-20180018540-A1) and Smith (United States Patent Application Publication US 2021/0049488 A1), as applied to claim 1 above, and further in view of Yang et. al. (Chinese Patent CN114549439 A).
Regarding claim 5, the combination of Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al. and Smith discloses the emotional engagement detection method based on positive emotional perception as claimed in claim 1. However, Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al. and Smith fail to disclose wherein the inputting the head pose direction and the smile intensity of the classroom student in the classroom video images and the access records of the learning system in the terminal system data for teaching resources, which are considered as information of complementary different modalities, into a deep network, to perform high-level semantic learning on the different modalities and perform feature fusion on the different modalities, to thereby obtain homogeneous fusion expression, specifically comprises: performing feature learning through asymmetric dual stream branches, wherein the asymmetric dual stream branches comprise a first branch and a second branch; the first branch is configured to utilize a deep cascaded autoencoder to perform feature mapping learning based on the head pose direction and the smile intensity of the classroom student in the classroom video images and the access records of the learning system through the deep network to achieve the high-level semantic learning; and the second branch is configured to extract low-level detail information for each of information of the different modalities; and connecting a generated higher-level feature of each of information of the different modalities to a lower-level feature of each of information of the different modalities through a skip layer, and generating a shared sparse fusion feature through an attention-guided feature cross fusion module, to correlate features of the different modalities together in a nonlinear manner and focus the features of the different modalities, and thereby obtain the homogeneous fusion expression.
Yang et. al. teaches wherein the inputting the head pose direction and the smile intensity of the classroom student in the classroom video images and the access records of the learning system in the terminal system data for teaching resources, which are considered as information of complementary different modalities, into a deep network, to perform high-level semantic learning on the different modalities and perform feature fusion on the different modalities, to thereby obtain homogeneous fusion expression, specifically comprises: performing feature learning through asymmetric dual stream branches, wherein the asymmetric dual stream branches comprise a first branch and a second branch; the first branch is configured to utilize a deep cascaded autoencoder to perform feature mapping learning based on the head pose direction and the smile intensity of the classroom student in the classroom video images and the access records of the learning system through the deep network to achieve the high-level semantic learning; and the second branch is configured to extract low-level detail information for each of information of the different modalities; and connecting a generated higher-level feature of each of information of the different modalities to a lower-level feature of each of information of the different modalities through a skip layer, and generating a shared sparse fusion feature through an attention-guided feature cross fusion module, to correlate features of the different modalities together in a nonlinear manner and focus the features of the different modalities, and thereby obtain the homogeneous fusion expression (Yang et. al., see specification paragraphs [0007]-[0036], discloses an RBG-D image semantic segmentation method based on multimodal feature fusion, an RGB-D image semantic segmentation method based on multimodal feature fusion).
These neural network features are critical to the claimed invention because it allows images to be inputted and subsequently categorized based on extracted sematic features. Yang et. al. describes the use of an input attention-guided multi-modal cross-fusion segmentation network model (ACFNe), which follows the encoder-decoder structure, with the ResNet-101 and ResNet-50 networks as backbone networks. All feature learning through asymmetric dual-stream branching, the high-level features of each information modality generated are connected by hop layers with features of their lower layers, by the attention-guided feature cross-fusion model, shared sparsely fused features are generated, and features of multiple modalities are correlated together in a complex non-linear manner to obtain a homogenous fused expression, using the fused augmented feature representation. Thus, it would have been obvious to one skilled in the art prior to the effective filing date of the claimed invention to have combined the teachings of Yang et. al. with the combination of Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al. and Smith to arrive at the solution presented by the claimed invention.
Claim(s) 6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhenzhen Luo et al., "Algorithm Modeling of Smile IntensityEstimation Using a Spatial Attention Convolutional Neural Network", "Journal of Physics:Conference Series", Vol. 1746 No. 1, 1-12, 20210131 in view of Liu et. al. (Chinese patent CN-116244474-A), Chen et. al. (Chinese patent CN-108399376-A), Wen et. al. (Chinese patent CN-112686195-A), Hazur et. al. (United States Patent Application Publication US-20180018540-A1) and Smith (United States Patent Application Publication US 2021/0049488 A1) as applied to claim 1 above, and further in view of Yu et. al. (Chinese Patent CN 108805089 B).
Regarding claim 6, the subject matter is substantially the same as claim 1 except for “wherein the deep learning network corresponding to the smile intensity estimation model comprises: a VGGNet and a ResNet; for the VGGNet, a pre-trained weight is loaded, the pre-trained weight is taken as a starting point for learning to compensate for a shortage of smile datasets, and to capture smile detailed information; and for the ResNet, a pre-trained weight is not loaded, an image feature is learned from scratch, and the ResNet is configured to focus on training information of the smile intensity; constructing a multi-task head pose estimation model by using a FSANet”. Thus, the rejection of claim 1 is incorporated herein.
Yu et. al. discloses wherein the deep learning network corresponding to the smile intensity estimation model comprises: a VGGNet and a ResNet; for the VGGNet, a pre-trained weight is loaded, the pre-trained weight is taken as a starting point for learning to compensate for a shortage of smile datasets, and to capture smile detailed information; and for the ResNet, a pre-trained weight is not loaded, an image feature is learned from scratch, and the ResNet is configured to focus on training information of the smile intensity; constructing a multi-task head pose estimation model by using a FSANet (Yang et. al. page 5, paragraphs 2-4: a multi-modal-based emotion recognition method based on facial image expression with fusion of VGG16 and RESNET 50 neural network architecture).
PNG
media_image2.png
764
1290
media_image2.png
Greyscale
This fusion neural network architecture is important to the claimed invention because it enables information extraction based on different objective functions, such as thinking activity of the student, the emotional state and cognitive attention of the student, which the neural network can train further for a subsequent learning task or used in the predictive model. The inventive concept is the multi-tasking head pose estimation model using FSANet networks, with each of the prior arts providing technical insights to the solution provided by the claimed invention. Thus, it would have been obvious to one skilled in the art prior to the effective filing date of the claimed invention to have combined the teachings of Zhenzhen et. al., Wen et. al., Hazur et. al., Liu et. al., Chen et. al., Smith and Yu et. al. to arrive at the emotional engagement detection method that is disclosed in the original application.
Conclusion
Response to Amendment
Examiner has carefully considered the amended claims and performed an updated search. New prior art was found to capture the features included in the amended claims.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JESSICA YIFANG LIN whose telephone number is (571)272-6435. The examiner can normally be reached M-F 7:00am-6:15pm, with optional day off.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached at 571-272-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JESSICA YIFANG LIN/Examiner, Art Unit 2668 August 6, 2026
/ANAND P BHATNAGAR/Primary Examiner, Art Unit 2668
August 6, 2026