Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Claims 1-4 and 6-20 are pending. Claims 1, 11, and 20 are independent.
This Application was published as US 20250232140.
Apparent priority is 17 January 2024.
The instant Application is directed to a method of translation based on sentiment analysis.
Applicant’s amendments and arguments are considered but are either unpersuasive or moot in view of the new grounds of rejection that, if presented, were necessitated by the amendments to the Claims.
This action is Final.
Response to Arguments
35 USC 101
Applicant's arguments have been fully considered and are persuasive. The rejection under 35 USC 101 is withdrawn.
35 USC 103
Applicant’s arguments with respect to 35 USC 103 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-4, 6-9, 11-18, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gupta et al. (US 20220358905 A1) in view of Poria et al. ("Fusing audio, visual and textual clues for sentiment analysis from multimodal content"), Cheng et al. (“Context-Aware Based Visual-Audio Feature Fusion for Emotion Recognition”), Wen et al. (CN 101593273 A), and Aher et al. (US 20220132217 A1).
Regarding claim 1, Gupta discloses: 1. A computing system comprising: a communications interface; ("[0191] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wire-line, optical fiber cable, radio frequency, etc., or any suitable combination of the foregoing…." )
a memory storing instructions; and ("[0189] The computer readable medium described in the claims below may be a computer readable signal medium or a computer readable storage medium..." )
at least one processor coupled to the communications interface and to the memory, the at least one processor being configured to execute the instructions to: ("[0192]... These computer program instructions may be provided to a processor of a general-purpose computer..." )
for a first content item, obtain text data, audio data, and video data, ("Input Transcription 108" ; "Input Audio 106" ; "Input Video 104" Fig. 1)
the text data and the audio data being associated with a first language, ("Translate input transcription from input language to output language 210" ; "Translate input audio from input language to output language 212" Fig. 2)
the first content item including a plurality of scenes; (“[0098]… For example, video preprocessor 124 may include facial landmark analysis, facial tracking algorithms, facial cropping and alignment algorithms, scene identification, and restoration and super resolution.”)
determine, by applying a first trained machine learning process to corresponding portions of the text data, the audio data and the video data of a first scene of the plurality of scenes, a sentiment score associated with the first scene, wherein the first trained machine learning process determines a respective sentiment value for each of the corresponding portions of the text data, ("[0064] In some embodiments, text preprocessor 128 is configured to convert text into phoneme analysis and/or perform emotional/sentiment analysis…" ; regarding machine learning, see also: “[0066] Emotion data includes any detectable emotion. Non-limiting examples of emotions include happy, sad, angry, scared, confused, excited, tired, sarcastic, disgusted, fearful, and surprised. Emotion data can further be compiled into a predetermined list of emotions and emotions can be communicated to the one or more processors and generators using computer readable formats, such as 1-hot or multi-class vectors, the second-to-last layer of a neural network or the output of a Siamese network to determine similarity. The same approach can be used for identifying and conveying the various other types of meta information.” – see also [0119], which discloses AI identifies meta information including emotion in each vocal segment, which would be within a scene.)
the audio data, and the video data and determines the sentiment score based on a combination of the respective sentiment values, and the sentiment value for the portions of the video data being based in part on a combination of environmental objects included in the first scene and a combination of colors associated with the first scene; (not explicitly disclosed)
generate, by applying a second trained machine learning process to the sentiment score and at least one of the text data and the audio data, translation data associated with the first scene, the translation data including a translation of the first scene and data indicating the sentiment score,
(“[0015]Once the meta data is acquired, the input transcription and input meta information are translated into the first output language based at least on the timing information and the emotion data, such that the translated transcription and meta information include similar emotion and pacing in comparison to the input transcription and input meta information…In some embodiments, translating the input transcription and input meta information includes providing the input transcription and input meta information to an AI transcription and meta translation generator configured to generate the translated transcription and meta information.” – generated meta information reads on data indicating the sentiment score; See also "[0125] Training AI meta information processor 130 to recognize and generate emotional data further improves the overall system because various sentiments can be captured and inserted into the translations…")
the translation being associated with a second language; and ("[0015] Once the meta data is acquired, the input transcription and input meta information are translated into the first output language” )
output the translation via at least one of a display device, an audio device or a combination thereof in accordance with the sentiment score and the translation, (Fig. 1, “Output Media File” 122)
the translation being output together with the data indicating the sentiment score. (Fig. 7 shows that the Translated Meta Info 135 and Translated Audio 140 are output together to the Video Sync Generator 144. [0076] discloses that the Translated Meta Info 135 can be provided to a user for review. Gupta does not explicitly disclose that the Translated Meta Info and Translated Audio are output to a display or audio device together.)
Gupta does not explicitly disclose that a sentiment score is based on a combination of scores for audio data and video data, or that the portions of video data identify and characterize a combination of environmental objects included in the first scene and a combination of colors associated with the first scene. Gupta also does not explicitly disclose that the translation is output to a display or audio device together with data indicating the sentiment score.
Poria discloses: determine, by applying a first trained machine learning process (“We have employed several supervised machine-learning-based classifiers for the sentiment classification task.” pg. 50, para 3)
to corresponding portions of the text data, the audio data and the video data of a first scene of the plurality of scenes, (“In YouTube dataset each video was segmented into several parts. According to the framerate of the video, we first converted each video segment into images. Then, for each video segment we extracted the facial features from all images and took the average to compute the final feature vector. Similarly, the audio and textual features were also extracted from each segment of the audio signal and text transcription of the video clip, respectively.” Pg. 53, section 4.4., first bullet point)
a sentiment score associated with the first scene, wherein the first trained machine learning process determines a respective sentiment value for each of the corresponding portions of the text data,
the audio data, and the video data and determines the sentiment score based on a combination of the respective sentiment values, ("Next, we fused the audio, visual and textual feature vectors to form a final feature vector which contained the information of both audio, visual and textual data. Later, a supervised classifier was employed on the fused feature vector to identify the overall polarity of each segment of the video clip. On the other hand, we also carried out an experiment on decision-level fusion, which took the sentiment classification result from 3 individual modalities as inputs and produced the final sentiment label as an output." Pg. 53, Section 4.4., second bullet point)
Gupta and Poria are considered analogous art to the claimed invention because they disclose methods of sentiment analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Gupta to use audio and video in combination with the text to determine sentiment as taught by Poria. Doing so would have been beneficial to enable effective extraction of the semantic and affective information conveyed during communication. (Poria, Pg. 51 Section 2.)
Poria does not explicitly disclose that the portions of video data identify and characterize a combination of environmental objects included in the first scene and a combination of colors associated with the first scene or that the translation is output to a display or audio device together with data indicating the sentiment score.
Cheng discloses: the sentiment value for the portions of the video data being based in part on a combination of environmental objects included in the first scene. (“Therefore, this paper proposes an emotion recognition method based on video scenes and objects context clues.” Pg. 1, last para; see also: “Due to scenes and objects in videos include abundant emotional clues, as shown in Fig.1, we propose a temporal-spatial network to exploit the scenes and objects information respectively. Specifically, we introduce hierarchical Bi-LSTM to summarize video scenes and combine the attention mechanism with the GCN to dig the emotion relationship between different objects." Pg. 2, section A. Overview. See also Fig. 5 which shows a video scene of a graveyard is identified as Sadness. Cheng discloses emotion relationship between different objects which implies a combination of objects. See also pg. 2, section C which discloses that the analysis is performed at a scene level.)
Gupta, Poria, and Cheng are considered analogous art to the claimed invention because they disclose methods of sentiment analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the system of Gupta in view of Poria to analyze environmental objects in the video data to determine sentiment, as disclosed by Cheng. Doing so would have been beneficial because scenes and objects in videos include abundant emotional clues (Cheng, Pg. 2, Section A) and because it would allow determining sentiment if there were not a person in the frame. Gupta discloses that emotional information improves translation (Gupta [0074]). This combination falls under combining prior art elements according to known methods to yield predictable results or simple substitution of one known element for another to obtain predictable results. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396.
Cheng does not explicitly disclose that the portions of video data identify and characterize a combination of colors associated with the first scene or that the translation is output to a display or audio device together with data indicating the sentiment score.
Wen discloses: A sentiment score based on a combination of colors associated with the first scene. (“In order to realize said purpose, the invention comprises the following steps: (1) converting the RGB color space into an HSL color space, to represent visual content in accordance with color space of human visual perception, (2) performing shot segmentation for video database, taking lens as the basic structure unit, further extracting lens low level characteristic vector, (3) the shot boundary detection to identify scene boundaries, the scene as a research unit, further extracting scene feature vector; (4) improved fuzzy comprehensive evaluation model, calculating the scene feature vector can reflect scene emotion information, (5) identifying basic emotion type generated by the scene audience using a high level feature vector and artificial neural network.” Pg. 2, para 3 – see also: “…then in each lens extracting key frame to represent the visual content of the lens, from the key frame extracting colour, texture, shape and so on, low level characteristic of the audio segment corresponding extracting lens at the same time, so as to obtain the shot or scene feature vector for emotion content analysis,…” pg. 1, last para – extracting the colors and using them as feature vectors for emotion content analysis reads on a sentiment score based on a combination of colors.)
Gupta, Poria, Cheng, and Wen are considered analogous art to the claimed invention because they disclose methods of sentiment analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination to further analyze colors in the video data as feature vectors to determine sentiment, as disclosed by Wen. Doing so would have been beneficial to improve identification accuracy (Wen pg. 2, first para). Gupta discloses that emotional information improves translation (Gupta [0074]). This combination falls under combining prior art elements according to known methods to yield predictable results or simple substitution of one known element for another to obtain predictable results. See MPEP 2141, KSR, 550 U.S. at 418, 82 USPQ2d at 1396. Poria discloses a fusion method in pg. 56, section 8.2, that one of ordinary skill in the art could have easily adapted to include two video decisions for both object and color based scores.
Wen does not explicitly disclose that the translation is output to a display or audio device together with data indicating the sentiment score.
Aher discloses: the translation data including a translation of the first scene and data indicating the sentiment score, (Fig. 3 shows emoticons for respective subtitles. An emoticon reads on data indicating a sentiment score. See also at least [0004] and [0007]-[0008] regarding selection of emoticons.)
And output the translation via at least one of a display device, an audio device or a combination thereof in accordance with the sentiment score and the translation, the translation being output together with the data indicating the sentiment score. (Fig. 1 shows a translation being output together with the emoticon (data indicating the sentiment score) on a display. See also: “[0004] … The media guidance application may then determine whether the identified keyword relates to an emotion corresponding to an emoticon by searching an emoticon database… In some embodiments, the media guidance application may cause the subtitles and the emoticon to be presented together. The emotion may be determined based on various factors, such as facial and body expressions, words in the dialogue, tone of the dialogue, and background music... Further, the media guidance application may then generate for display, at the subtitle region or another location of the media asset's video frame, the first subtitle data including the determined emoticon. By inserting emoticons based on the keywords and other factors, the media guidance application may more precisely convey the emotion than conventional systems that rely on simple translations in the subtitle data.”)
Gupta, Poria, Cheng, Wen, and Aher are considered analogous art to the claimed invention because they disclose methods of sentiment analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination to further output an emoticon with the translation as disclosed by Aher. Doing so would have been beneficial to more precisely convey the emotion (Aher [0004]).
Regarding claim 2, Gupta, Poria, Cheng, and Wen do not disclose the additional limitations.
Aher discloses: 2. The computing system of claim 1, wherein a portion of the audio data is associated with music associated with the first scene, and wherein the first trained machine learning process determines the respective sentiment value for the corresponding portion of the audio data by performing music analysis on the music, the respective sentiment value indicating, along a range of sentiment values, a degree to which the music corresponds to a sentiment, the sentiment being a tone or an emotion. ("[0013]… In some embodiments, the media guidance application may calculate an emotion score to determine if an emoticon is necessary to enhance the viewing experience. For example, the media guidance application may consider a facial expression of an actor of the media asset; a body movement of an actor in the media asset; words in a dialogue in the audio portion of the media asset; a tone of the dialogue in the audio portion of the media asset; and background music in the audio portion of the media asset. Each of the considerations has a weighted value to arrive at a total emotion score. …")
Gupta, Poria, Cheng, Wen, and Aher are considered analogous art to the claimed invention because they disclose methods of sentiment analysis. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination to further analyze background music as disclosed by Aher. Doing so would have been beneficial to more precisely target the emotion (Aher [0007]) and so a comprehensive weighted score could be calculated (Aher [0089]).
Regarding claim 3, Gupta discloses: 3. The computing system of claim 1, wherein the respective sentiment value for the corresponding portion of audio data is determined by performing waveform analysis on the corresponding portion of the audio data. ("[0050] Audio preprocessor 126 may include processes for partitioning audio content for each speaker into separate audio tracks, removing or cleaning up background noise, and enhancing voice quality data. These processes may be performed using any known systems and methods capable of performing the processes enumerated herein. In some embodiments, audio preprocessor 126 is also configured to automatically identify input language 110 using voice recognition software such as those known in the art..” preprocessing the audio before determining Meta Info reads on performing waveform analysis of the audio.)
Regarding claim 4, Gupta disclose: 4. The computing system of claim 1, wherein the respective sentiment value for the corresponding portion of text data is based on a combination of words included in the corresponding portion of the text data. (“[0115] In some embodiments of the present invention, text preprocessor 128 is a preprocessing AI. AI text preprocessor 128 may include processes for detecting and analyzing phonemes within text such as input transcription 108. AI text preprocessor 128 may further include processes for detecting and analyzing emotions/sentiments within text, parts of speech, proper nouns, and idioms. These processes may be performed using any known AI preprocessors capable of performing the processes enumerated herein. For example, AI text preprocessor 128 may include phonetic analysis based in IPA or similar system generated through dictionary lookup or transformer model or GAN model, sentiment analysis, parts of speech analysis, proper noun analysis, and idiom detection algorithms.” – idioms reads on a combination of words)
Regarding claim 6, Gupta discloses: 6. The computing system of claim 1, wherein the translation of the first scene is translated text data associated with the second language. ("Translated Transcription 134" Fig. 5)
Regarding claim 7, Gupta discloses: 7. The computing system of claim 1, wherein the translation of the first scene is translated audio data associated with the second language. ("Translated Audio 140" Fig. 6)
Regarding claim 8, Gupta, Poria, Cheng, and Wen do not disclose the additional limitations.
Aher discloses: 8. The computing system of claim 1, wherein a portion of the audio data is associated with music associated with the first scene, and the respective sentiment value for the corresponding portion of the audio data is based at least in part on the music. ("[0013]… In some embodiments, the media guidance application may calculate an emotion score to determine if an emoticon is necessary to enhance the viewing experience. For example, the media guidance application may consider a facial expression of an actor of the media asset; a body movement of an actor in the media asset; words in a dialogue in the audio portion of the media asset; a tone of the dialogue in the audio portion of the media asset; and background music in the audio portion of the media asset…")
See claim 2 for motivation statement.
Regarding claim 9, Gupta discloses: 9. The computing system of claim 1, wherein the audio data is associated with dialogue associated with the first scene. ("[0055] Speaker diarization processor 125 is configured to partition input audio 106 into homogeneous vocal segments according to an identifiable speaker. Ultimately, speaker diarization processor 125 performs a series of steps to identify one or more speakers in input media 102 and associate each string of speech (also referred to as a vocal segment) with the proper speaker." - voice segments are dialogue.)
Claim 11 is a method claim with limitations corresponding to the limitations of Claim 1 and is rejected under similar rationale.
Claim 12 is a method claim with limitations corresponding to the limitations of Claim 2 and is rejected under similar rationale.
Claim 13 is a method claim with limitations corresponding to the limitations of Claim 3 and is rejected under similar rationale.
Claim 14 is a method claim with limitations corresponding to the limitations of Claim 4 and is rejected under similar rationale.
Claim 15 is a method claim with limitations corresponding to the limitations of Claim 6 and is rejected under similar rationale.
Claim 16 is a method claim with limitations corresponding to the limitations of Claim 7 and is rejected under similar rationale.
Claim 17 is a method claim with limitations corresponding to the limitations of Claim 8 and is rejected under similar rationale.
Claim 18 is a method claim with limitations corresponding to the limitations of Claim 9 and is rejected under similar rationale.
Claim 20 is a computer readable medium claim with limitations corresponding to the limitations of Claim 1 and is rejected under similar rationale.
Claim(s) 10 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Gupta in view of Poria, Cheng, Wen, and Aher as applied in claim 1 above, in further view of Kofman et al. (US 20210165973 A1).
Regarding claim 10, Gupta, Poria, Cheng, Wen, and Aher do not disclose the additional limitations.
Kofman discloses: 10. The computing system of claim 1, wherein the at least one processor is further configured to: receive a search query; search the text data; search the translated data; and return a search result. ("[0200] As can be seen in FIG. 13, in some example embodiments, the transcribed language UI 1300A includes a search field(s) 1360, which can be used to quickly find specified text in the viewed transcribed language and translated language transcripts..." – the transcribed language is the original language (text data).)
Gupta, Poria, Cheng, Wen, Aher, and Kofman are considered analogous art to the claimed invention because they are in the field of speech processing. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the combination with a search engine to search both original and translated texts as taught by Kofman. Doing so would have been beneficial so that people around the world can receive diverse content (Kofman [0004]), and so that the user could search in either language.
Claim 19 is a method claim with limitations corresponding to the limitations of Claim 10 and is rejected under similar rationale.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Knight et al. (US 9788777 B1). Knight discloses emotion identification in music (see col 9, last para – col 10 first para) by waveform analysis (see col 16, second para and also Fig. 3).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JON C MEIS whose telephone number is (703)756-1566. The examiner can normally be reached Monday - Thursday, 8:30 am - 5:30 pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Hai Phan can be reached at 571-272-6338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JON CHRISTOPHER MEIS/Examiner, Art Unit 2654
/HAI PHAN/Supervisory Patent Examiner, Art Unit 2654