DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) was submitted on 04/25/2024. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Specification
Applicant is reminded of the proper language and format for an abstract of the disclosure.
The abstract should be in narrative form and generally limited to a single paragraph on a separate sheet within the range of 50 to 150 words in length. The abstract should describe the disclosure sufficiently to assist readers in deciding whether there is a need for consulting the full patent text for details.
The language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc. In addition, the form and legal phraseology often used in patent claims, such as “means” and “said,” should be avoided.
Drawings
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference character(s) not mentioned in the description: 6b in Figure 2, 86b in Figure 5, 72 in figure 8. Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Objections
Claim 33-35, 38 is objected to because of the following informalities:
Claim 33 details the limitation “recording calibration data of the user by use of the recording device as input information in the main storage unit or the temporary storing unit” which should read “recording calibration data of the user by use of the recording device as input information in a main storage unit or a temporary storing unit” as this is the first mention of “main storage unit” and “temporary storage unit” as Claim 33 is an independent Claim and “a” should be used as the antecedence rather than “the” for the first mention of “main storage unit” and “temporary storage unit”.
Claims 33 and 38 details the limitation “analyzing the input data by the processing unit by comparing input data to the reference data” which should read as “analyzing the input data by the processing unit by comparing input data to a reference data”. Claim 33 is an independent Claim and ‘a” should be used as the antecedence basis as this is the first mention of “reference data”. Claim 38 is ultimately dependent on Claim 20, which but Claim 20 does not mention “reference data”, thus the reference data in Claim 38 is the first mention of reference data in the limitation, thus the antecedence should be “a” rather than “the”.
Claim 34 details the limitations “comprising the steps of - preparing calibration data
- preparing input data, wherein the calibration data and the input data is video data of the user and/or audio data of the user, wherein
the preparing includes at least one of
- extracting bounding boxes to detect frontal faces, and
- cropping video data and
- resizing video data and
- splitting audio data into time windows of a certain length and
- splitting audio data into frames.”
Examiner suggests that the claim limitation consistently use commas after each element in the list and with only using the conjunction “and” after the next to last element in this list, such that it the claim limitation would read as:
“comprising the steps of:
- preparing calibration data,
- preparing input data,
wherein the calibration data and the input data is video data of the user and/or audio data of the user, wherein the preparing includes at least one of:
- extracting bounding boxes to detect frontal faces,
- cropping video data,
- resizing video data,
- splitting audio data into time windows of a certain length, and
- splitting audio data into frames.”
Claim 35 details the limitation “wherein an emotional feature extracted is at least one of - a fundamental frequency of the voice - a formant frequency of the voice - jitter of the voice - shimmer of the voice - intensity of the voice” which should read “wherein an emotional feature extracted is at least one of - a fundamental frequency of the voice, - a formant frequency of the voice, - jitter of the voice, - shimmer of the voice, or - intensity of the voice”. That is, the list should include commas and the conjunction “or” between the next to last element and the last element of the list.
Claim 38 details “A computer program product comprising instructions which when the program is executed by a computer processing unit of a detection device of claim 20cause the computer to carry out method steps for detecting an emotional state of a user, comprising the steps of…” which should read “A computer program product comprising instructions which when the program is executed by a computer processing unit of a detection device of claim 20 to cause the computer to carry out method steps for detecting an emotional state of a user, comprising the steps of…”
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 26 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 26 details the limitation “wherein the connection device comprises a wired or wireless connection element”. As claim 26 is dependent on Claim 20, and Claim 20 details “a connection element”, it is not clear whether the “connection device” of Claim 26 is intended to be the same as the “connection element” of claim 20 or a different device. Examiner interprets the “connection device” as the “connection element” of Claim 20.
Claims 33-38 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 33 and 38 detail the limitation “providing at least one of a detection device and a system, a recording device and an interface device”. As the limitation details “at least one of” then details a list consisting of “a detection device and a system, a recording device and an interface device” with commas and multiple conjunctions (“ands”), it is not clear nor distinct what the claimed subject matter is detailing as there can be multiple interpretations. As the limitation details:
Detection device
System
Recording device
Interface device
It is not clear nor distinct whether the claim limitation requires at least one of (a) + (b) and at least one of (c) + (d) or at least one (a) + at least one of (b) + at least one of (c) + at least one of (d).
Examiner interprets the limitation as at least one of (a) + (b) and at least one of (c) + (d).
Claims 34-37 are rejected due to dependence on Claim 33.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 20-21,23-24, 26-28, 31, 33-34, 36, and 38 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Srivastava (US20190090020).
In regards to Claim 20, Srivastava (US20190090020) teaches “a processing unit for processing data (control circuitry includes processor 204 – [0062]),
- a main data storage unit for storing data (memory with local database – [0062]),
- a connecting element for connecting the detection device to an interface device (transceiver with communication via a communication network – [0062]),
wherein the detection device is adapted to be calibrated to said user by use of the processing unit and calibration data (calibration system configured to predict a set of new highlight points and lowlight points and in the post-calibration stage the calibrated set of markers are utilized to map the emotions expressed by the audience in real time or near-real time – [0060]-[0061]), and wherein the processing unit is adapted to analyze input data based on said calibration (“In the post calibration stage, the calibrated set of markers may be utilized to precisely map the emotions expressed by the audience 112 in real time or near-real time. A measure of an impact of different marked scenes may be analyzed based on the set of common peaks in the plurality of different types of input signals. Such impact measure may be utilized to derive a rating for each marked scene of the media item 104 and a cumulative rating for the entire media item. Further, calibration system 102 may be configured to generate, using the simulation engine, a video score, an audio score, and a distribution score for the media item 104 based on an accuracy of the capture of emotion response data from the audience 112 . Such scores (such as the video score, the audio score, and the distribution score) may be utilized to compare a first media item (e.g., a first movie), with a second media item (e.g. a second movie) of the set of media items. A final rating or impact of different media items may be identified based on the comparison of scores of different media items with each other” – [0061]; “The processor 204 may be configured normalize different types of input signals of different users in the audience 112 , captured by the plurality of different types of sensors 114 . Such normalization may be done based on a degree of expressive ability of different users of the audience 112 . The processor 204 may be further configured to detect an emotional response from each user in the audience 112 based on the plurality of input video signals. Examples of the emotional response may include, but are not limited to, happy, sad, surprise, anger, fear, disgust, contempt, and neutral emotional response. The processor 204 may be further configured to measure, using the face recognizer 210 , an emotional response level of the detected emotional response of different users of the audience 112 in the first time interval. The emotional response level of each user of the audience 112 may indicate a degree by which a user in the audience 112 may express an emotional response in the first time interval” – [0084]; “For example, the processor 204 may be configured to measure, using the face recognizer 210 , the emotional response level of a smile as a response from a first user of the audience 112 . The processor 204 may be further configured to assign, using the face recognizer 210 , the emotional response level on a scale from “0” to “1” for the smile response from the first user. Similarly, in some instances, when a user shows a grin face, the processor 204 may assign, using the face recognizer 210 , an average value for the emotional response level, i.e. a value of “0.5” for the user that shows a grin face. In other instances, when another user laughs, the processor 204 may assign, using the face recognizer 210 , a maximum value of “1” for the user that laughs. The face recognizer 210 may be configured to determine the emotional response level of each of the audience 112 at a plurality of time instants within the first time interval” – [0085]).”
In regards to Claim 21, Srivastava discloses the claimed invention as detailed above. Srivastava further teaches “wherein the processing unit is adapted to process said input data and said calibration data by preparing the input data and the calibration data and extracting emotions features of said input data and said calibration data (“The calibration system 102 may be configured to receive the plurality of different types of input signals from the plurality of different types of sensors 114 . The plurality of different types of input signals may include audio signals from the set of audio sensors 114 A, video signals that includes face data of the audience 112 from the set of image sensors 114 B, and biometric signals from the set of biometric sensors 114 C. Examples of audio signals may include, but are not limited to, audio data of claps, laughter, chatter, and whistles of the audience 112 . Example of the face data may include but is not limited to a sequence of facial images (or a video feed) of the audience members. The sequence of facial images of the audience members may comprise specific expressions of the audience members. The face data may further comprise eye-tracking, motion-tracking, and gesture tracking information associated with the audience members. Examples of the biometric signals may include, but are not limited to, pulse rate data, heart rate data, body heat data, and body temperature data of each member of the audience 112 . Such plurality of different types of input signals may correspond to a set of emotional responses of the audience 112 at the playback of the media item 104 at the media device 110 . Thereafter, the calibration system 102 may be configured to generate an amalgamated audience response signal based on the synchronized plurality of different input signals and a plurality of weights assigned to the plurality of different types of sensors 114 that captures the plurality of different types of input signals” – [0051]; “Different members of the audience 112 may exhibit a different degree of expressive ability, which may be captured in the plurality of different types of input signals in real time or near-real time from the plurality of different types of sensors 114 . For example, at a certain comic scene at the playback of the media item 104 , different members of the audience 112 may exhibit different types of smiles or laughter. In order to robustly extract an aggregate emotional response for the comic scene, a normalization of the plurality of different types of input signals may be done. Therefore, the calibration system 102 may be configured to normalize the plurality of different types of input signals for each scene of the media item 104 based on different parameters of the audience 112” – [0052]; “accordance with an embodiment, the calibration system 102 may be configured to compute a reaction delay for a first emotional response of the audience 112 captured by the plurality of different types of sensors 114 for a first scene of the media item 104 with respect to a position of the first marker 106 . The first marker 106 may span a first time slot of the set of time slots in the media item 104 . The first time slot may correspond to the first scene in the media item 104 . Similar to computation of the reaction delay for the first scene, the calibration system 102 may be configured to compute a set of reaction delays for a set of emotional responses of the audience 112 . Such set of reaction delays may be captured by the plurality of different types of sensors 114 for a set of scenes in the media item 104 with respect to a position of the set of markers. Each marker of the set of markers may span one time slot of the set of time slots in the media item 104” – [0055]).”
In regards to Claim 23, Srivastava discloses the claimed invention as detailed above. Srivastava further teaches “wherein the processing unit is adapted to calculate an emotional state based on multiple input data and multiple comparisons of said input data to said calibration data (“Each of the plurality of scenes may include the expected-emotions-tagging-metadata (for example, one or more emotional tags 302 , such as neutral, happy, sad, and anger), along the timeline 300 of the media item. The expected-emotions-tagging metadata may indicate an associative relationship between the set of time slots and a set of specified emotional states that may be expected from the audience 112 at the set of time slots at playback of the media item. Examples of the set of specified emotional states may include, but is not limited to happy, neutral, sad, and anger” – [0115]; “At 1254 , a set of common positive peaks and a set of common negative peaks may be identified in each of the plurality of different types of input signals. The processor 204 may be configured to identify the set of common positive peaks and the set of common negative peaks based on the overlay of the plurality of different types of input signals” – [0172]; “At 1256 , a plurality of highpoints and a plurality of lowlight points for a plurality of scenes of the media item 104 may be calculated. The processor 204 may be configured to calculate the plurality of highlight points and the plurality of lowlight points based on the identified set of common positive peaks and the identified set of common negative peaks” – [0173]; “At 1268 , a first set of highlight points and a second set of lowlight points in the media item 104 may be predicted. The processor 204 may be configured to predict the first set of highlight points and the second set of lowlight points based on a change in control parameters using a simulation engine, wherein the control parameters include a geographical region, a race, a physical dimension of recording room, an age, a gender, a genre of the media item 104 . The control passes to end 1270” – [0179]).”
In regards to Claim 24, Srivastava discloses the claimed invention as detailed above. Srivastava further teaches “wherein the detection device comprises a self-supervised learning computer program structure for learning to extract meaningful features of data of the user for emotion detection, the self-supervised learning computer program structure is adapted to learn to predict matching audio and video data of the user (“It may be noted that the technique described for the simulation and estimation of the emotional response and calibration of media items for different audiences from different geographies and demographics may not be so limited. Thus, the aforementioned emotional response may be estimated through any suitable technique, such as an auto encoder for a new set of geography data, classifiers, extractors for emotional response data, neural network-based regression and boosted decision tree regression for emotional response data, and supervised machine learning models or deep learning models to precisely estimate value-based weight values” – [0113]).”
In regards to Claim 26, Srivastava discloses the claimed invention as detailed above. Srivastava further teaches “wherein the connection device comprises a wired or wireless connection element (“The transceiver 214 may comprise suitable logic, circuitry, and interfaces that may be configured to communicate with other electronic devices, via the communication network 120 . The transceiver 214 may implement known technologies to support wireless communication. The transceiver 214 may include, but are not limited to an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, and/or a local buffer circuitry. The transceiver 214 may communicate via offline and online wireless communication with networks, such as the Internet, an Intranet, and/or a wireless network, such as a cellular telephone network, a wireless local area network (WLAN), personal area network, and/or a metropolitan area network (MAN). The wireless communication may use any of a plurality of communication standards, protocols and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), wideband code division multiple access (W-CDMA), code division multiple access (CDMA), LTE, time division multiple access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (such as IEEE 802.11, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n, and/or any other IEEE 802.11 protocol), voice over Internet Protocol (VoIP), Wi-MAX, Internet-of-Things (IoT) technology, Machine-Type-Communication (MTC) technology, a protocol for email, instant messaging, and/or Short Message Service (SMS).” – [0068]).”
In regards to Claim 27, Srivastava discloses the claimed invention as detailed above of “A system for detecting the emotional state of a user comprising a detection device according to claim 20”. Srivastava further teaches “at least one recording device for sensing the user (“Thereafter, the processor 204 may be configured to initially pre-calibrate the plurality of different types of sensors 114 prior to the playback of the media item (such as the media item 104 ), for which emotional response data is to be captured from the audience 112 The pre-calibration of the set of audio sensors 114 A, the set of image sensors 114 B and the set of biometric sensors 114 C of the plurality of different types of sensors 114 may ensure that credible data from the plurality of different types of sensors 114 is captured at next processing stage.” – [0071]).”
In regards to Claim 28, Srivastava discloses the claimed invention as detailed above. Srivastava further teaches “wherein the system is adapted to be calibrated to the user by use of the processing unit and calibration data recorded by the recording device (“In operation, control signals may be received at the calibration system 102 to initialize acquisition of emotional response data based on different sensor-based data and audience 112 -based data. Such emotional response data may be utilized to precisely update durations of different high points/low points in a media item (for example, the media item 104 ), which may be scheduled for selected/scheduled for playback at the media device 110 . Such control signals may be routed to a processor 204 , such as the processor 204 of the calibration system 102 . Thereafter, the processor 204 may be configured to initially pre-calibrate the plurality of different types of sensors 114 prior to the playback of the media item (such as the media item 104 ), for which emotional response data is to be captured from the audience 112 The pre-calibration of the set of audio sensors 114 A, the set of image sensors 114 B and the set of biometric sensors 114 C of the plurality of different types of sensors 114 may ensure that credible data from the plurality of different types of sensors 114 is captured at next processing stage” – [0071]).”
In regards to Claim 31, Srivastava discloses the claimed invention as detailed above. Srivastava further teaches “wherein the system comprises an interface device for interaction between the system and the user, the interface device being connected to the detection device by the connecting element (“The calibration system 102 may comprise suitable logic, circuitry, and interfaces that may be configured to capture an audience response at playback of the media item 104 . The calibration system 102 may be further configured to analyze the audience response to calibrate the plurality of markers in the media item 104 of the set of media items. Alternatively stated, the calibration system 102 may embed one or more markers in the media item 104 and calibrate a position of the plurality of markers in the media item 104 based on an audience response at specific durations of playback of the media item 104 . In some embodiments, the calibration system 102 may be implemented as an internet of things (IOT) enabled system that operates locally in the network environment 100 or remotely, through the communication network 120 . In some embodiments, the calibration system 102 may be implemented as a standalone device that may be integrated within the media device 110 . Examples of implementation of the calibration system 102 may include a projector, a smart television (TV), a personal computer, a special-purpose device, a media receiver, such as a set top box (STB), a digital media player, a micro-console, a game console, an (High Definition Multimedia Interface) HDMI compliant source device, a smartphone, a tablet computer, a personal computer, a laptop computer, a media processing system, or a calibration device” – [0036]).”
In regards to Claim 33, Srivastava teaches “- providing at least one of a detection device and a system, a recording device and an interface device (“The calibration system 102 may comprise suitable logic, circuitry, and interfaces that may be configured to capture an audience response at playback of the media item 104 . The calibration system 102 may be further configured to analyze the audience response to calibrate the plurality of markers in the media item 104 of the set of media items. Alternatively stated, the calibration system 102 may embed one or more markers in the media item 104 and calibrate a position of the plurality of markers in the media item 104 based on an audience response at specific durations of playback of the media item 104 . In some embodiments, the calibration system 102 may be implemented as an internet of things (IOT) enabled system that operates locally in the network environment 100 or remotely, through the communication network 120 . In some embodiments, the calibration system 102 may be implemented as a standalone device that may be integrated within the media device 110 . Examples of implementation of the calibration system 102 may include a projector, a smart television (TV), a personal computer, a special-purpose device, a media receiver, such as a set top box (STB), a digital media player, a micro-console, a game console, an (High Definition Multimedia Interface) HDMI compliant source device, a smartphone, a tablet computer, a personal computer, a laptop computer, a media processing system, or a calibration device” – [0036]; “The calibration system 102 may be further configured to pre-calibrate the set of biometric sensors 114 C of the plurality of different types of sensors 114 before the playback of the media item 104 by the media device 110 . Such pre-calibration of the set of biometric sensors 114 C may be done based on a measurement of biometric data in the test environment in presence of the test audience 116 at the playback of the test media. A standard deviation in the measured biometric data of the test audience 116 may be determined at the playback of the test media. The calibration system 102 may be further configured to assign a weight to each sensor of the plurality of different types of sensors 114 based on different estimations done for different types of sensors. For example, the set of audio sensors 114 A may be weighted based on different levels of the captured noise components. The set of image sensors 114 B may be weighted based on error rate in detection of number of faces. The set of biometric sensors 114 C may be weighted based on deviations in the measured biometric data of the test audience 116 . The detailed operation of pre-calibration has been discussed in detail in FIG. 2. The pre-calibration of the plurality of different types of sensors 114 enables obtaining of credible data from the plurality of different types of sensors 114” – [0046]; calibration system includes processor, memory, and transceiver with communication via a communication network – [0062]),
- recording calibration data of the user by use of the recording device as input information in the main storage unit or the temporary storing unit (“The calibration system 102 may store the expected-emotions-tagging metadata as a pre-stored data in a database managed by the calibration system 102 . In some embodiments, the calibration system 102 may fetch the expected-emotions-tagging metadata from a production media server (or dedicated content delivery networks) and store the expected-emotions-tagging metadata in the memory of the calibration system 102 . In certain scenarios, the calibration system 102 may be configured to store the expected-emotions-tagging metadata in the media item 104 . The expected-emotions-tagging metadata may be a data structure with a defined schema, such as a scene index number (I), a start timestamp (T0 ), an end timestamp (T 1 ), and an expected emotion (EM). In certain scenarios, the expected-emotions-tagging metadata may comprise an expected emotion type and an expected sub-emotion type. An example of the stored expected-emotions-tagging metadata associated with the media item 104 (for example, a movie with a set of scenes), is given below, for example, in Table 1.” – [0049]; memory with local database – [0062]),
- calibrating the detection device to the user by processing the calibration data by a processing unit of the detection device (calibration system configured to predict a set of new highlight points and lowlight points and in the post-calibration stage the calibrated set of markers are utilized to map the emotions expressed by the audience in real time or near-real time – [0060]-[0061]; calibration unit includes control circuitry which includes processor 204 – [0062]),
- recording input data by the recording device (memory with local database – [0062]; “the processor 204 may be configured to receive the plurality of different types of input signals from the plurality of different types of sensors 114 . The plurality of different types of input signals may include audio signals from the set of audio sensors 114 A, video signals that includes face data of the audience 112 from the set of image sensors 114 B, and biometric signals from the set of biometric sensors 114 C.” – [0082]),
- analyzing the input data by the processing unit by comparing input data to the reference data (“Alternatively stated, the set of highlight points may be identified based on how evenly spread the points are and what were the individual values for audio, video and pulse recorded compared to the actual emotional response from the media item 104 .” – [0103]; “The processor 204 may be configured to generate the set of new highlight points for different types of audiences based on predictive analysis and comparison of the media item 104 and the emotional response with a set of scores estimated for the media item 104 . For example, the set of scores may include an audio score, a video score and a distribution score. The processor 204 may be configured to assign a video score, an audio score, and a distribution score to the media item 104 based on the captured plurality of different types of input signals captured by the plurality of different types of sensors 114 at the playback of the media item 104 . The processor 204 may be configured to compute the audio score based on the audio signals in the plurality of different types of input signals” - [0111]),
- determining a commonality of the input data and the reference data (“In the post calibration stage, the calibrated set of markers may be utilized to precisely map the emotions expressed by the audience 112 in real time or near-real time. A measure of an impact of different marked scenes may be analyzed based on the set of common peaks in the plurality of different types of input signals. Such impact measure may be utilized to derive a rating for each marked scene of the media item 104 and a cumulative rating for the entire media item. Further, calibration system 102 may be configured to generate, using the simulation engine, a video score, an audio score, and a distribution score for the media item 104 based on an accuracy of the capture of emotion response data from the audience 112 . Such scores (such as the video score, the audio score, and the distribution score) may be utilized to compare a first media item (e.g., a first movie), with a second media item (e.g. a second movie) of the set of media items. A final rating or impact of different media items may be identified based on the comparison of scores of different media items with each other” – [0061]; “The processor 204 may be configured normalize different types of input signals of different users in the audience 112 , captured by the plurality of different types of sensors 114 . Such normalization may be done based on a degree of expressive ability of different users of the audience 112 . The processor 204 may be further configured to detect an emotional response from each user in the audience 112 based on the plurality of input video signals. Examples of the emotional response may include, but are not limited to, happy, sad, surprise, anger, fear, disgust, contempt, and neutral emotional response. The processor 204 may be further configured to measure, using the face recognizer 210 , an emotional response level of the detected emotional response of different users of the audience 112 in the first time interval. The emotional response level of each user of the audience 112 may indicate a degree by which a user in the audience 112 may express an emotional response in the first time interval” – [0084]; “For example, the processor 204 may be configured to measure, using the face recognizer 210 , the emotional response level of a smile as a response from a first user of the audience 112 . The processor 204 may be further configured to assign, using the face recognizer 210 , the emotional response level on a scale from “0” to “1” for the smile response from the first user. Similarly, in some instances, when a user shows a grin face, the processor 204 may assign, using the face recognizer 210 , an average value for the emotional response level, i.e. a value of “0.5” for the user that shows a grin face. In other instances, when another user laughs, the processor 204 may assign, using the face recognizer 210 , a maximum value of “1” for the user that laughs. The face recognizer 210 may be configured to determine the emotional response level of each of the audience 112 at a plurality of time instants within the first time interval” – [0085]).”
In regards to Claim 34, Srivastava discloses the claimed invention as detailed above. Srivastava further teaches “comprising the steps of
– preparing calibration data - preparing input data, wherein the calibration data and the input data is video data of the user and/or audio data of the user (“the processor 204 may be configured to receive the plurality of different types of input signals from the plurality of different types of sensors 114 . The plurality of different types of input signals may include audio signals from the set of audio sensors 114 A, video signals that includes face data of the audience 112 from the set of image sensors 114 B, and biometric signals from the set of biometric sensors 114 C.” – [0082]), wherein
the preparing includes at least one of - extracting bounding boxes to detect frontal faces, and - cropping video data and - resizing video data and - splitting audio data into time windows of a certain length and - splitting audio data into frames (“The face recognizer 210 may comprise suitable logic, circuitry, and interfaces that may be configured to detect a change in facial expression of different users in the audience 112 by the set of image sensors 114 B of the plurality of different types of sensors 114 . The face recognizer 210 may be further configured to detect static facial expressions of different users in the audience 112 by the set of image sensors 114 B of the plurality of different types of sensors 114 . The face recognizer 210 may be implemented as a software application or a hardware circuit, such as an Application-Specific Integrated Circuit (ASIC) processor” – [0066]; “The video frame clipper 212 may comprise suitable logic, circuitry, and interfaces that may be configured to generate a plurality of highpoint media segments based on the identified set of highlight points and the set of lowlight points. The video frame clipper 212 may be implemented based on the processor 204 , such as one of a programmable logic controller (PLC), a microcontroller, an X86-based processor, a Reduced Instruction Set Computing (RISC) processor, an Application-Specific Integrated Circuit (ASIC) processor, a Complex Instruction Set Computing (CISC) processor, and/or other hardware processors” – [0067]; “At the media playback stage, for different granular durations of the playback (for example, 60 seconds) of the media item 104 , the processor 204 may be configured to receive the plurality of different types of input signals from the plurality of different types of sensors 114 . The plurality of different types of input signals may include audio signals from the set of audio sensors 114 A, video signals that includes face data of the audience 112 from the set of image sensors 114 B, and biometric signals from the set of biometric sensors 114 C. Such plurality of different types of input signals may be stored in the memory 216 as an array of a specific size in accordance with a specific duration of the playback, for example, a signal array with “60” values may include values for “60 seconds” of input signal, where “60” denotes a maximum index value of the signal array that stores a value for a type of input signal at every second at an index position within the signal array” – [0082]).”
In regards to Claim 36, Srivastava discloses the claimed invention as detailed above. Srivastava further teaches “comprising the steps of - recording one or multiple, input data of the user (memory with local database – [0062]; “the processor 204 may be configured to receive the plurality of different types of input signals from the plurality of different types of sensors 114 . The plurality of different types of input signals may include audio signals from the set of audio sensors 114 A, video signals that includes face data of the audience 112 from the set of image sensors 114 B, and biometric signals from the set of biometric sensors 114 C.” – [0082]),- determining an emotional state of the user (“The processor 204 may be configured normalize different types of input signals of different users in the audience 112 , captured by the plurality of different types of sensors 114 . Such normalization may be done based on a degree of expressive ability of different users of the audience 112 . The processor 204 may be further configured to detect an emotional response from each user in the audience 112 based on the plurality of input video signals. Examples of the emotional response may include, but are not limited to, happy, sad, surprise, anger, fear, disgust, contempt, and neutral emotional response. The processor 204 may be further configured to measure, using the face recognizer 210 , an emotional response level of the detected emotional response of different users of the audience 112 in the first time interval. The emotional response level of each user of the audience 112 may indicate a degree by which a user in the audience 112 may express an emotional response in the first time interval” – [0084]; “For example, the processor 204 may be configured to measure, using the face recognizer 210 , the emotional response level of a smile as a response from a first user of the audience 112 . The processor 204 may be further configured to assign, using the face recognizer 210 , the emotional response level on a scale from “0” to “1” for the smile response from the first user. Similarly, in some instances, when a user shows a grin face, the processor 204 may assign, using the face recognizer 210 , an average value for the emotional response level, i.e. a value of “0.5” for the user that shows a grin face. In other instances, when another user laughs, the processor 204 may assign, using the face recognizer 210 , a maximum value of “1” for the user that laughs. The face recognizer 210 may be configured to determine the emotional response level of each of the audience 112 at a plurality of time instants within the first time interval” – [0085]), - outputting the determined emotional state on the interface device (“The calibration system 102 may comprise suitable logic, circuitry, and interfaces that may be configured to capture an audience response at playback of the media item 104 . The calibration system 102 may be further configured to analyze the audience response to calibrate the plurality of markers in the media item 104 of the set of media items. Alternatively stated, the calibration system 102 may embed one or more markers in the media item 104 and calibrate a position of the plurality of markers in the media item 104 based on an audience response at specific durations of playback of the media item 104 . In some embodiments, the calibration system 102 may be implemented as an internet of things (IOT) enabled system that operates locally in the network environment 100 or remotely, through the communication network 120 . In some embodiments, the calibration system 102 may be implemented as a standalone device that may be integrated within the media device 110 . Examples of implementation of the calibration system 102 may include a projector, a smart television (TV), a personal computer, a special-purpose device, a media receiver, such as a set top box (STB), a digital media player, a micro-console, a game console, an (High Definition Multimedia Interface) HDMI compliant source device, a smartphone, a tablet computer, a personal computer, a laptop computer, a media processing system, or a calibration device” – [0036]; “The amalgamated audience response signal may be associated with one or more emotional states of a set of defined emotional states. The fourth user interface 1000 further comprises an application card 1004 to display the set of defined emotional states such as happiness, sadness and disgust. The fourth user interface 1000 may be utilized as a dashboard interface to visualize the amalgamated emotional response data of different audiences by different users, such as content owners, content creators/producers, or content distributors. The fourth user interface 1000 may display a plurality of highlight points 1006 of the amalgamated audience response signal represented in the first graph 1002 . The first graph 1002 may be a dynamic moving chart which represents a value of the amalgamated audience response signal at a particular timestamp (of the media item 104 ), when the respective timestamp of the media item 104 may be played. The fourth user interface 1000 may enable a user to fast forward and/or seek a plurality of timestamps of the media item 104 . In cases where the user fast forwards and/or seek seeks to a particular timestamp of the media item 104 , the first graph 1002 may represent the amalgamated audience response signal of the audience 112 at the respective timestamp of the media item 104” – [0141]).”
In regards to Claim 38, Srivastava discloses the claimed invention as detailed above in Claim 20. Srivastava further teaches “providing at least one of a detection device and a system, a recording device and an interface device (“The calibration system 102 may comprise suitable logic, circuitry, and interfaces that may be configured to capture an audience response at playback of the media item 104 . The calibration system 102 may be further configured to analyze the audience response to calibrate the plurality of markers in the media item 104 of the set of media items. Alternatively stated, the calibration system 102 may embed one or more markers in the media item 104 and calibrate a position of the plurality of markers in the media item 104 based on an audience response at specific durations of playback of the media item 104 . In some embodiments, the calibration system 102 may be implemented as an internet of things (IOT) enabled system that operates locally in the network environment 100 or remotely, through the communication network 120 . In some embodiments, the calibration system 102 may be implemented as a standalone device that may be integrated within the media device 110 . Examples of implementation of the calibration system 102 may include a projector, a smart television (TV), a personal computer, a special-purpose device, a media receiver, such as a set top box (STB), a digital media player, a micro-console, a game console, an (High Definition Multimedia Interface) HDMI compliant source device, a smartphone, a tablet computer, a personal computer, a laptop computer, a media processing system, or a calibration device” – [0036]; “The calibration system 102 may be further configured to pre-calibrate the set of biometric sensors 114 C of the plurality of different types of sensors 114 before the playback of the media item 104 by the media device 110 . Such pre-calibration of the set of biometric sensors 114 C may be done based on a measurement of biometric data in the test environment in presence of the test audience 116 at the playback of the test media. A standard deviation in the measured biometric data of the test audience 116 may be determined at the playback of the test media. The calibration system 102 may be further configured to assign a weight to each sensor of the plurality of different types of sensors 114 based on different estimations done for different types of sensors. For example, the set of audio sensors 114 A may be weighted based on different levels of the captured noise components. The set of image sensors 114 B may be weighted based on error rate in detection of number of faces. The set of biometric sensors 114 C may be weighted based on deviations in the measured biometric data of the test audience 116 . The detailed operation of pre-calibration has been discussed in detail in FIG. 2. The pre-calibration of the plurality of different types of sensors 114 enables obtaining of credible data from the plurality of different types of sensors 114” – [0046]; calibration system includes processor, memory, and transceiver with communication via a communication network – [0062]),
- recording calibration data of the user by use of the recording device as input information in the main storage unit or the temporary storing unit (“The calibration system 102 may store the expected-emotions-tagging metadata as a pre-stored data in a database managed by the calibration system 102 . In some embodiments, the calibration system 102 may fetch the expected-emotions-tagging metadata from a production media server (or dedicated content delivery networks) and store the expected-emotions-tagging metadata in the memory of the calibration system 102 . In certain scenarios, the calibration system 102 may be configured to store the expected-emotions-tagging metadata in the media item 104 . The expected-emotions-tagging metadata may be a data structure with a defined schema, such as a scene index number (I), a start timestamp (T0 ), an end timestamp (T 1 ), and an expected emotion (EM). In certain scenarios, the expected-emotions-tagging metadata may comprise an expected emotion type and an expected sub-emotion type. An example of the stored expected-emotions-tagging metadata associated with the media item 104 (for example, a movie with a set of scenes), is given below, for example, in Table 1.” – [0049]; memory with local database – [0062]),
- calibrating the detection device to the user by processing the calibration data by a processing unit of the detection device (calibration system configured to predict a set of new highlight points and lowlight points and in the post-calibration stage the calibrated set of markers are utilized to map the emotions expressed by the audience in real time or near-real time – [0060]-[0061]; calibration unit includes control circuitry which includes processor 204 – [0062]),
- recording input data by the recording device (memory with local database – [0062]; “the processor 204 may be configured to receive the plurality of different types of input signals from the plurality of different types of sensors 114 . The plurality of different types of input signals may include audio signals from the set of audio sensors 114 A, video signals that includes face data of the audience 112 from the set of image sensors 114 B, and biometric signals from the set of biometric sensors 114 C.” – [0082]),
- analyzing the input data by the processing unit by comparing input data to the reference data (“Alternatively stated, the set of highlight points may be identified based on how evenly spread the points are and what were the individual values for audio, video and pulse recorded compared to the actual emotional response from the media item 104 .” – [0103]; “The processor 204 may be configured to generate the set of new highlight points for different types of audiences based on predictive analysis and comparison of the media item 104 and the emotional response with a set of scores estimated for the media item 104 . For example, the set of scores may include an audio score, a video score and a distribution score. The processor 204 may be configured to assign a video score, an audio score, and a distribution score to the media item 104 based on the captured plurality of different types of input signals captured by the plurality of different types of sensors 114 at the playback of the media item 104 . The processor 204 may be configured to compute the audio score based on the audio signals in the plurality of different types of input signals” - [0111]),
- determining a commonality of the input data and the reference data (“In the post calibration stage, the calibrated set of markers may be utilized to precisely map the emotions expressed by the audience 112 in real time or near-real time. A measure of an impact of different marked scenes may be analyzed based on the set of common peaks in the plurality of different types of input signals. Such impact measure may be utilized to derive a rating for each marked scene of the media item 104 and a cumulative rating for the entire media item. Further, calibration system 102 may be configured to generate, using the simulation engine, a video score, an audio score, and a distribution score for the media item 104 based on an accuracy of the capture of emotion response data from the audience 112 . Such scores (such as the video score, the audio score, and the distribution score) may be utilized to compare a first media item (e.g., a first movie), with a second media item (e.g. a second movie) of the set of media items. A final rating or impact of different media items may be identified based on the comparison of scores of different media items with each other” – [0061]; “The processor 204 may be configured normalize different types of input signals of different users in the audience 112 , captured by the plurality of different types of sensors 114 . Such normalization may be done based on a degree of expressive ability of different users of the audience 112 . The processor 204 may be further configured to detect an emotional response from each user in the audience 112 based on the plurality of input video signals. Examples of the emotional response may include, but are not limited to, happy, sad, surprise, anger, fear, disgust, contempt, and neutral emotional response. The processor 204 may be further configured to measure, using the face recognizer 210 , an emotional response level of the detected emotional response of different users of the audience 112 in the first time interval. The emotional response level of each user of the audience 112 may indicate a degree by which a user in the audience 112 may express an emotional response in the first time interval” – [0084]; “For example, the processor 204 may be configured to measure, using the face recognizer 210 , the emotional response level of a smile as a response from a first user of the audience 112 . The processor 204 may be further configured to assign, using the face recognizer 210 , the emotional response level on a scale from “0” to “1” for the smile response from the first user. Similarly, in some instances, when a user shows a grin face, the processor 204 may assign, using the face recognizer 210 , an average value for the emotional response level, i.e. a value of “0.5” for the user that shows a grin face. In other instances, when another user laughs, the processor 204 may assign, using the face recognizer 210 , a maximum value of “1” for the user that laughs. The face recognizer 210 may be configured to determine the emotional response level of each of the audience 112 at a plurality of time instants within the first time interval” – [0085]).”
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 22 and 25 are rejected under 35 U.S.C. 103 as being unpatentable over Srivastava in view of Bosnak (US20230021339).
In regards to Claim 22, Srivastava discloses the claimed invention as detailed above in Claim 21. Srivastava further teaches “wherein the processing unit comprises a deep neural network (“the aforementioned emotional response may be estimated through any suitable technique, such as an auto encoder for a new set of geography data, classifiers, extractors for emotional response data, neural network-based regression and boosted decision tree regression for emotional response data, and supervised machine learning models or deep learning models to precisely estimate value-based weight values” – [0113]).”
Srivastava is silent with regards to the language of “wherein the processing unit is adapted to embed the extracted emotions features of the input data and the calibration data into the deep neural network and to create an emotional landscape of the user”
Bosnak teaches “wherein the processing unit is adapted to embed the extracted emotions features of the input data and the calibration data into the deep neural network and to create an emotional landscape of the user (“In some embodiments, a particular emotional label may be applied to the given interaction based on the biomarker measurements, e.g., using the Associated Dimensional Affect emotion wheel or other suitable classification model or any combination thereof. Peaks in arousal may create differentiated, and often conflicting, emotional states over a longer section of time (e.g. 20 seconds or more). The progression and variation of emotional states over time may form a signature morphology of emotions that can be called an Embodied State. The system for the avatar may be configured to continuously search memory of previously stored interactions with similar signatures to the given interaction performed by the user. When an embodied state appears more frequently than a defined threshold, the software may be configured to label the state as expression of a Complex. When a Complex is detected, the system may refer back to previous moments when the Complex was active and generate a response asking the user for connections. The re-emergence of the same Complex may unconsciously generate a similar embodied ANS response. As a result, the system may record Complexes, embodied states and user responses to learn the paradoxical emotional landscape of the user creating a library of Complex responses. The more memories that are generated, the greater the sophistication and extent of the Complex library for improved recognition of the user's embodied state at any given interaction. For example, the system may be increasingly aware of emotional regularities in the user and can compare the verbal communication during those Complex moments. The system may remind the user of the similarity between those moments. As a result, the avatar may be generated based on the user's emotional Complexes to simulate strong empathy with the user, thus eliciting in the User a sense of being understood.” – [0056]; using neural networks – [0098])”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Srivastava to incorporate the teaching of Bosnak to utilize the data collected to create an emotional landscape of the user. By utilizing data to create an emotional landscape this is an improvement that yields predictable results in the identification and determination of the emotional state of a user.
In regards to Claim 25, Srivastava discloses the claimed invention as detailed above in Claim 20. Srivastava is silent with regards to the language of “wherein at least one of the main data storage unit and an additional temporary storage unit is adapted for temporarily storing the input information.”
Bosnak teaches “wherein at least one of the main data storage unit and an additional temporary storage unit is adapted for temporarily storing the input information (“In some embodiments, the attuned avatar system 100 may include, e.g., a storage device 101 . In some embodiments, the data storage solution of the storage device 101 may include, e.g., a suitable memory or storage solutions for maintaining electronic data representing the activity histories for each account. For example, the data storage solution may include database technology such as, e.g., a centralized or distributed database, cloud storage platform, decentralized system, server or server system, among other storage systems. In some embodiments, the data storage solution may, additionally or alternatively, include one or more data storage devices such as, e.g., a hard drive, solid-state drive, flash drive, or other suitable storage device. In some embodiments, the data storage solution may, additionally or alternatively, include one or more temporary storage devices such as, e.g., a random-access memory, cache, buffer, or other suitable memory device, or any other data storage solution and combinations thereof” – [0068]).”
It would have been obvious to one of ordinary skill in the art to modify Srivastava to incorporate the teaching of Bosnak to utilize a computer that has two memories – one that does permanent storage and one that does temporary storage. By utilizing a computer that includes permanent and temporary storage this is an improvement that is well known in the art that yields predictable results in the operation of a computing system as computer systems are well known to have permanent storage devices (hard drives, SSDs) and temporary storage devices (DRAM).
Claims 29-30 and 32 are rejected under 35 U.S.C. 103 as being unpatentable over Srivastava in view of Goldstein (US20200236473)
In regards to Claim 29, Srivastava discloses the claimed invention as detailed above in Claim 27. Srivastava further teaches “wherein the system comprises a piece, wherein the recording device is part of this piece (“In some embodiments, the set of biometric sensors 114 C may be non-invasively attached to the body of each member of the audience 112” – [0041]).”
Srivastava is silent with regards to the language of “wherein the system comprises an ear-piece, wherein the recording device is part of this ear-piece”
Goldstein teaches “wherein the system comprises an ear-piece, wherein the recording device is part of this ear-piece (“The diagram assumes that the earpiece includes a full complement of functions including always on recording, biometric measuring and recording, sound pressure level measurements from both an ambient microphone and an ear canal microphone, voice activity detection, key word detection and analysis, personal audio assistant functions, transmission of data to a phone or a server or cloud device, among many other functions” – [0022])”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Srivastava to incorporate the teaching of Goldstein to utilize an earpiece to be non-invasively attached to the body of a user. By utilizing an earpiece this is an improvement that yields predictable results when monitoring the audio data with respect to a user.
In regards to Claim 30, Srivastava in view of Goldstein discloses the claimed invention as detailed above. Srivastava further teaches “wherein the piece comprises at least one or more additional recording devices (“The set of biometric sensors 114 C may be a plurality of sensors that may be configured to capture biometric data of the audience 112 at the playback of the media item 104 of the set of media items. Examples of a biometric sensor may include but is not limited to, a pulse rate sensor, a breath rate sensor, a body temperature sensor, or a skin conductance sensor, or other specialized sensors to measure different emotions aroused in the audience 112 at the playback of the set of media items. In some embodiments, the set of biometric sensors 114 C may be non-invasively attached to the body of each member of the audience 112 .” – [0041]).”
Srivastava is silent with regards to the language of “wherein the ear-piece comprises at least one of a speaker and one or more additional recording devices”
Goldstein further teaches “wherein the ear-piece comprises at least one of a speaker and one or more additional recording devices (“The diagram assumes that the earpiece includes a full complement of functions including always on recording, biometric measuring and recording, sound pressure level measurements from both an ambient microphone and an ear canal microphone, voice activity detection, key word detection and analysis, personal audio assistant functions, transmission of data to a phone or a server or cloud device, among many other functions” – [0022])”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Srivastava to incorporate the teaching of Goldstein to utilize an earpiece to be non-invasively attached to the body of a user. By utilizing an earpiece this is an improvement that yields predictable results when monitoring the audio data with respect to a user.
In regards to Claim 32, Srivastava discloses the claimed invention as detailed above in Claim 27. Srivastava further teaches “wherein the system comprises multiple recording devices, wherein the recording devices comprise at least one of a temperature sensor (“Examples of a biometric sensor may include but is not limited to, a pulse rate sensor, a breath rate sensor, a body temperature sensor, or a skin conductance sensor, or other specialized sensors to measure different emotions aroused in the audience 112 at the playback of the set of media items. In some embodiments, the set of biometric sensors 114 C may be non-invasively attached to the body of each member of the audience 112 . In other embodiments, the set of biometric sensors 114 C may be invasively implanted in the body of each member of the audience 112 , for example, an ingestible capsule with nano-biometric sensors. In some implementations, instead of a set of biosensors, a single standalone sensor device may be installed in the closed environment, to collectively detect a heart signature and further monitor the biometric data of all the members of the audience 112 .” – [0041]).”
Srivastava is silent with regards to the language of “wherein the system comprises multiple recording devices, wherein the recording devices comprise at least one of an acceleration sensor and a temperature sensor and a humidity sensor”
Goldstein teaches “wherein the system comprises multiple recording devices, wherein the recording devices comprise at least one of an acceleration sensor and a temperature sensor and a humidity sensor(“ Exemplary physiological and environmental sensors that may be incorporated into a Bluetooth® or other type of earpiece module include, but are not limited to accelerometers, auscultatory sensors, pressure sensors, humidity sensors, color sensors, light intensity sensors, pulse oximetry sensors, pressure sensors, and neurological sensors, etc” – [0134]; “The sensors can constitute biometric, physiological, environmental, acoustical, or neurological among other classes of sensors. In some embodiments, the sensors can be embedded or formed on or within an expandable element or balloon or other material that is used to occlude (or partially occlude) the ear canal. Such sensors can include non-invasive contactless sensors that have electrodes for EEGs, ECGs, transdermal sensors, temperature sensors, transducers, microphones, optical sensors, motion sensors or other biometric, neurological, or physiological sensors that can monitor brainwaves, heartbeats, breathing rates, vascular signatures, pulse oximetry, blood flow, skin resistance, glucose levels, and temperature among many other parameters. The sensor(s) can also be environmental including, but not limited to, ambient microphones, temperature sensors, humidity sensors, barometric pressure sensors, radiation sensors, volatile chemical sensors, particle detection sensors, or other chemical sensors. The sensors can be directly coupled to a processor or wirelessly coupled via a wireless communication system. Also note that many of the components can be wirelessly coupled (or coupled via wire) to each other and not necessarily limited to a particular type of connection or coupling” – [0135])”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Srivastava to incorporate the teaching of Goldstein to utilize an earpiece to be non-invasively attached to the body of a user and to have a plurality of different sensors associated with the earpiece to collect data on the user. By utilizing an earpiece this is an improvement that yields predictable results when monitoring the data with respect to a user.
Claim 35 is rejected under 35 U.S.C. 103 as being unpatentable over Srivastava in view of Hossain (Nazia Hossain et. al., “Finding Emotion from Multi-Lingual Voice Data”, July 13-17th, 2020, 2020 IEEE 44th Annual computers, Software, and Applications Conference, https://ieeexplore.ieee.org/document/9202764)
In regards to Claim 35, Srivastava discloses the claimed invention as detailed above. Srivastava further teaches “comprising the steps of - extracting emotional features of the calibration data,- extracting emotional features of the input data (“The pre-calibration may be done to precisely capture emotional response data from the audience 112 at the playback of the media item 104 , with a minimal effect of noise signals or measurement errors of different sensors on the captured emotional response data” – [0044]; “Different members of the audience 112 may exhibit a different degree of expressive ability, which may be captured in the plurality of different types of input signals in real time or near-real time from the plurality of different types of sensors 114 . For example, at a certain comic scene at the playback of the media item 104 , different members of the audience 112 may exhibit different types of smiles or laughter. In order to robustly extract an aggregate emotional response for the comic scene, a normalization of the plurality of different types of input signals may be done. Therefore, the calibration system 102 may be configured to normalize the plurality of different types of input signals for each scene of the media item 104 based on different parameters of the audience 112” – [0052]; “In accordance with an embodiment, the normalized plurality of different type of input signals may be utilized to determine a peak emotional response level for each emotional response from each user of the audience 112 . Thereafter, the normalized plurality of different types of input signals may be synchronized and overlaid in a timeline that may be same as a playback timeline of the media item 104 . A set of common positive peaks and a set of common negative peaks in each of the plurality of different types of input signals may be further identified based on the overlay of the plurality of different types of input signals. An example of the set of common positive peaks and the set of common negative peaks is shown in FIG. 6B and FIG. 11. The calibration system 102 may be configured to calculate a plurality of highlight points and a plurality of lowlight points for a plurality of scenes of the media item 104 . Such plurality of highlight points or the plurality of lowlight points may be calculated based on the identified set of common positive peaks and the set of common negative peaks. In certain scenarios, the calibration system 102 may be configured to identify the set of common positive points and the set of common negative points without using the expected-emotion-tagging metadata associated with the media item 104 . The set of common positive points and the set of common negative points may be identified based on the set of emotional responses of the audience 112 at the playback of the set of scenes of the media item 104” – [0058]; “In accordance with an embodiment, the plurality of different types of input signals may be processed for emotional response estimation from each type of input signal followed by an estimation of a reaction delay of the audience 112 for a specific scene of the media item 104 . In accordance with another embodiment, the plurality of different types of input signals may be processed for an emotional response estimation from each type of input signal followed by an estimation of periods of highlight points and low light points in the media item 104” – [0083]; “For example, the processor 204 may be configured to measure, using the face recognizer 210 , the emotional response level of a smile as a response from a first user of the audience 112 . The processor 204 may be further configured to assign, using the face recognizer 210 , the emotional response level on a scale from “0” to “1” for the smile response from the first user. Similarly, in some instances, when a user shows a grin face, the processor 204 may assign, using the face recognizer 210 , an average value for the emotional response level, i.e. a value of “0.5” for the user that shows a grin face. In other instances, when another user laughs, the processor 204 may assign, using the face recognizer 210 , a maximum value of “1” for the user that laughs. The face recognizer 210 may be configured to determine the emotional response level of each of the audience 112 at a plurality of time instants within the first time interval” - [0085]).”
Srivastava is silent with regards to the language of “wherein an emotional feature extracted is at least one of - a fundamental frequency of the voice - a formant frequency of the voice - jitter of the voice - shimmer of the voice - intensity of the voice”
Hossain teaches “wherein an emotional feature extracted is at least one of - a fundamental frequency of the voice - a formant frequency of the voice - jitter of the voice - shimmer of the voice - intensity of the voice (“Therefore in our research, detecting continuous and qual itative speech has being focused. The continuous speech features like pitch, intensity, energy, and jitter values are extracted from the voice signals. These acoustic characteristics of speech carry linguistic contents of emotional states. Murray et al. showed comparative changes among different voice features in different emotions” – Page 410).”
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Srivastava to incorporate the teaching of Hossain to extract emotional features from audio data. By focusing on speech features including pitch, intensity, energy, and jitter this is an improvement to the extraction of emotional states from audio data of a suer.
Allowable Subject Matter
Claim 37 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims.
In regards to Claim 37, Srivastava discloses the claimed invention as detailed above. Srivastava is silent with regards to the language of “wherein the method comprises the steps of - providing a Siamese neural network,- feeding in a first batch of matching audio and video data of multiple subjects, the audio and video data are separated from each other, into the Siamese neural network, - predicting a correlation between said audio data and said video data of the first batch, - feeding in a second batch of matching video and audio data separately of a single subject concerning different emotional states of the subject into the Siamese neural network, - predicting a correlation between said audio data and said video data of the second batch.”
Examiner’s Note
The following prior art is considered to be of interest to the case:
West (US20170310927) details a method for presenting emotion animation in an audio or video call in which the system can visually identify the emotional state of the user by sampling various facial points of the user.
Tao (CN113255800A) details an emotion modeling system based on audio and video data.
Chappell (US20230047787) teaches a system for controlling progress of audio-video content based on sensor data of multiple users, composite neuro-physiological state and content engagement power.
Newton (US20120041917) teaches a method of adapting an environment of a terminal that receives data from a user in an environment associated with the mental state of a user.
Li (CN112911334A) teaches an emotion recognition method based on audio and video data.
Gordon (US20190075239) teaches an method of storing video content until a change in an emotional or cognitive state of a user.
Feinauer (US20200176018) teaches a computing platform to extract audio and video segments and utilizing machine learning to produce probability distribution of emotional states of a person.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YOSSEF KORANG-BEHESHTI whose telephone number is (571)272-3291. The examiner can normally be reached Monday - Friday 10:00 am - 6:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Catherine Rastovski can be reached at (571) 270-0349. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YOSSEF KORANG-BEHESHTI/ Primary Examiner, Art Unit 2857