Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The disclosure is objected to because of the following informalities:
“unititioned” in paragraph [0071] appears to be a typo
Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-2 and 8-10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Mowzoon, US Pub. 2011/0093263 A1 (hereinafter Mowzoon) in view of Parthasarathi et al., US Pub. 10798271 B2 (hereinafter Parthasarathi), and further in view of Skarbovsky et al., US Pub. 2018/0143956 A1 (hereinafter Skarbovsky).
In regards to Claim 1, Mowzoon discloses an artificial intelligence-based subtitle management apparatus comprising: a subtitle generation unit configured to acquire content data including video data and audio data from a first user terminal (Mowzoon: [Abstract], “An automated closed captioning, captioning, or subtitle generation system that automatically generates the captioning text from the audio signal in a submitted online video”), and to generate first subtitle data synchronized based on time information of the audio data using a subtitle generation model (Mowzoon: Fig. 3 and [0022], where the web site utilizes the current speech recognition model to generate the text transcript from the audio portion of the data…The corrected file is added to the original file to generate a texted file. Text file gets added back as caption layer for use by the video; [0020], where this information combined with the accompanied timestamp can then be fed into any number of the captioning solutions); and a subtitle modification unit configured to receive a subtitle modification request including modification data from the second user terminal (Mowzoon: [Abstract], where the user text review and correction step allows the text prediction model to accumulate additional corrected data with each use), and to generate second subtitle data by modifying the first subtitle data based on the modification data (Mowzoon: [0022], where the corrected file is added to the original file to generate a texted file). However, Mowzoon fails to explicitly disclose a content provision unit configured to re-synchronize the first subtitle data based on motion information of the video data, and to provide the content data and the first subtitle data to a second user terminal in a matched manner. He also fails to explicitly disclose wherein the video data comprises a plurality of sequentially continuous frames, wherein the content provision unit is further configured to: classify the plurality of frames into a plurality of groups based on the motion information, wherein, when an N-th frame (where N is a positive integer) and an (N+1)-th frame among the plurality of frames include different motion information, the N-th and (N+1)-th frames are classified into different groups; and wherein the subtitle modification unit is further configured to: determine a suitability of the subtitle modification request based on at least one of a matching rate between the modification data and the first subtitle data, and information related to the subtitle modification requester; and when the suitability exceeds a threshold, modify the first subtitle data to generate the second subtitle data.
Parthasarathi from a similar endeavor teaches a content provision unit configured to re-synchronize the first subtitle data based on motion information of the video data (Parthasarathi: [Abstract], “the subtitle timing application determines that a temporal edge associated with a subtitle does not satisfy a timing guideline based on a shot change. The shot change occurs within a sequence of frames of an audiovisual program. The subtitle timing application then determines a new temporal edge that satisfies the timing guideline relative to the shot change. Subsequently, the subtitle timing application causes a modification to a temporal location of the subtitle within the sequence of frames based on the new temporal edge.”), and to provide the content data and the first subtitle data to a second user terminal in a matched manner (Parthasarathi: Fig. 4 and Col. 14-15, lines 66-67 and 1, where at step 422 , the subtitle GUI generator 290 generates the subtitle GUI 190 and causes the subtitle GUI 190 to be displayed to the user); and wherein the video data comprises a plurality of sequentially continuous frames (Parthasarathi: [Abstract], where the shot change occurs within a sequence of frames of an audiovisual program), wherein the content provision unit is further configured to: classify the plurality of frames into a plurality of groups based on the motion information (Parthasarathi: Col. 4, lines 16-20, where the visual component 132 includes any number of different shot sequences (not shown), where each shot sequence includes a set of frames that usually have similar spatial-temporal properties and run for an uninterrupted period of time), wherein, when an N-th frame (where N is a positive integer) and an (N+1)-th frame among the plurality of frames include different motion information, the N-th and (N+1)-th frames are classified into different groups (Parthasarathi: Col. 4, lines 21-26, where the frame at which one shot sequence ends and a different shot sequence begins is referred to herein as a shot change 152. As a general matter, each frame included in the visual component 132 is related to a particular time during the playback of the audiovisual program 130 via the frame rate 136).
Skarbovsky from a similar endeavor teaches that the subtitle modification unit is further configured to: determine a suitability of the subtitle modification request based on at least one of a matching rate between the modification data and the first subtitle data, and information related to the subtitle modification requester (Skarbovsky: [0050], where an aggregation engine 160 may set an audience threshold so that at least x% of the audience devices 150 must suggest the same edit before it is approved…a suggested edit must meet a given confidence threshold for the phonemes of the word (s) that it is to replace before the audience threshold is met. Audience devices 150 that are associated with a sufficient number of good edits for a given event, or that are designated as trusted editors, may be given greater weight than audience devices 150 not associated with a sufficient number of good edits); and when the suitability exceeds a threshold, modify the first subtitle data to generate the second subtitle data (Skarbovsky: [0050], where an aggregation engine 160 may set an audience threshold so that at least x% of the audience devices 150 must suggest the same edit before it is approved).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Mowzoon in view of Parthasarathi and Skarbovsky such that the artificial intelligence-based subtitle management apparatus disclosed by Mowzoon would also include the limitations further disclosed by Parthasarathi and Skarbovsky. Mowzoon supplies the base architecture, Parthasarathi explains the re-synching and frame-grouping, and Skarbovsky describes the technique for filtering subtitle edits based on suitability thresholds. This combination allows for a clearer and more streamlined self-correcting captioning pipeline.
Regarding Claim 2, the combined teaching of Mowzoon, Parthasarathi, and Skarbovsky discloses the apparatus of claim 1, wherein the plurality of groups comprises a first group and a second group that are sequentially continuous, wherein the content provision unit is further configured to: when a portion of the first subtitle data corresponding to the first group also corresponds to motion information of the second group, synchronize a start point of the portion corresponding to the first group to match a start time of a first frame of the second group (Parthasarathi: Col. 1-2, lines 66-67, 1-9, where the method includes performing one or more operations, via a processor, to determine that a first temporal edge associated with a first subtitle does not satisfy a timing guideline relative to a first s hot change that occurs within a sequence of frames of an audiovisual program; performing one or more additional operations, via the processor, to select a second temporal edge that satisfies the timing guideline relative to the first shot change; and causing a temporal location of the first subtitle within the sequence of frames to be modified based on the second temporal edge).
Regarding Claim 8, the combined teaching of Mowzoon, Parthasarathi, and Skarbovsky discloses the apparatus of claim 1, wherein the subtitle modification unit is further configured to: determine a higher suitability when the subtitle modification requester is a content provider than when the requester is a content viewer, and determine a higher suitability as the subtitle modification history, similar content viewing history, or content provision history of the requester increases (Skarbovsky: [0050], where audience devices 150 that are associated with a sufficient number of good edits for a given event, or that are designated as trusted editors, may be given greater weight than audience devices 150 not associated with a sufficient number of good edits (or are associated with a sufficient number of bad edits).
Regarding Claim 9, the combined teaching of Mowzoon, Parthasarathi, and Skarbovsky discloses a method for managing subtitles using artificial intelligence, the method comprising all of the claimed limitations (see rejection of claim 1 for details).
Regarding Claim 10, the combined teaching of Mozoon, Parthasarathi, and Skarbovsky discloses a non-transitory computer-readable recording medium storing a program that, when executed by a computer, causes the computer to perform the method of claim 9 (Parthasarathi: Col. 20, lines 64-65, where there are one or more non-transitory computer-readable storage media including instructions; Skarbovsky: [0007], where examples are implemented as a computer process, a computing system, or as an article of manufacture such as device, computer program product, or computer readable medium. According to an aspect, the computer program product is a computer storage medium readable by a computer system and encoding a computer program comprising instructions for executing a computer process).
Claim(s) 3 is/are rejected under 35 U.S.C. 103 as being unpatentable over Mowzoon in view of Parthasarathi and Skarbovsky, and further in view of Barreira Avegliano et al., US Pub. 9609397 B1 (hereinafter Barreira Avegliano).
Regarding Claim 3, the combined teaching of Mowzoon, Parthasarathi, and Skarbovsky discloses the apparatus of claim 1, but fails to explicitly disclose wherein the content provision unit is further configured to: when a start point of the first subtitle data and a start point of the second subtitle data corresponding to a modified portion differ based on the subtitle modification request, synchronize the start point of the second subtitle data with the start point of the first subtitle data.
Barreira Avegliano from a similar endeavor teaches synchronizing the start point of a second subtitle data with the start point of a first subtitle data when the start point of the first subtitle data and the start point of the second subtitle data corresponding to a modified portion differ based on the subtitle modification request (Barreria Avegliano: Col. 21, lines 23-27 and lines 37-48, where audio fingerprints may be used in a manner that guarantees the real-time adjustment of subtitles whose timestamps are not correctly aligned with the audio/video playback times, thereby allowing perfect synchronization of the medias… (a) media player starts the execution of multimedia stream and parsing of the captioning file (that contains the fingerprints); (b) the media player seeks to the times of the first and the last anchors and calculates the audio fingerprint of the media file at these specific moments; (c) media player compares the audio fingerprints from the captioning file against the ones extracted from the multimedia stream. If they match, the files are considered synchronized. If not, the media player seeks to the beginning of the multimedia stream and computes all of its fingerprints until a match with the first anchor is found). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Mowzoon, Parthasarathi, and Skarbovsky in view of Barreria Avegliano such that their disclosed apparatus of claim 1 would also include Barreria Avegliano’s method of synchronizing the start points of the subtitle data. This would ensure that the subtitles correctly correspond to their designated media file portion.
Claim(s) 4-6 is/are rejected under 35 U.S.C. 103 as being unpatentable over Mowzoon in view of Parthasarathi, Skarbovsky, and Barreira Avegliano, and further in view of Yegnanarayanan, US Pub. 2013/0035961 A1 (hereinafter Yegnanarayanan).
In regards to Claim 4, the combined teaching of Mowzoon, Parthasarathi, Skarbovsky, and Barreira Avegliano discloses the apparatus of claim 3, but fails to explicitly disclose wherein the subtitle modification unit is further configured to: when the second subtitle data is generated, extract a modification keyword from the modification data, and update the second subtitle data by modifying a portion including the modification keyword to reflect the same.
Yegnanarayanan from a similar endeavor teaches extracting a modification keyword from the modification data when the second subtitle data is generated, and updating the second subtitle data by modifying a portion including the modification keyword to reflect the same (Yegnanarayanan: [0159], where when the user makes a correction to the initial set of medical facts automatically extracted from the text narrative, fact review component 106 and/or fact extraction component 104 may learn from the user's correction and may apply it to other texts and/or to other portions of the same text…the system may then search for one or more other portions of the text that are similar, e.g., because they share the same features (or a suitable subset of those features) as the text portion corresponding to the fact identified by the user…When a match is found, in some embodiments the system may automatically extract, from the text portion with the matching features, a fact similar to the fact identified by the user for the first portion of the text). The feature that the system extracts from the user’s correction correlates to the claim’s modification keyword, and the updating of the matching text portions correlates to modifying a portion including the modification keyword. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Mowzoon, Parthasarathi, Skarbovsky, and Barreria Avegliano in view of Yegnanarayanan such that their disclosed apparatus of claim 3 would also include Yegnanarayanan’s method of extracting a modification keyword from the modification data when the second subtitle data is generated, and updating the second subtitle data by modifying a portion including the modification keyword to reflect the same. This would help the system correct and line up the right words in the subtitles.
Regarding Claim 5, the combined teaching of Mowzoon, Parthasarathi, Skarbovsky, Barreira Avegliano, and Yegnanarayanan discloses the apparatus of claim 4, wherein the subtitle modification unit is further configured to: search for similar content data related to the modification keyword among other content data for which subtitle data has been generated by the subtitle generation unit, and modify a portion including the modification keyword in the subtitle data of the similar content data in the same manner (Yegnanarayanan: [0159], where when a match is found, in some embodiments the system may automatically extract, from the text portion with the matching features, a fact similar to the fact identified by the user for the first portion of the text; [0160], where in response to the user correction, the system may generate one or more rules specifying that a fact of the type corresponding to the user-identified fact should be extracted from text that is similar to the text portion corresponding to the user's correction (e.g., text sharing features with the text portion corresponding to the user's correction)).
Regarding Claim 6 the combined teaching of Mowzoon, Parthasarathi, Skarbovsky, Barreira Avegliano, and Yegnanarayanan discloses the apparatus of claim 5, wherein the subtitle modification unit is further configured to provide the modification data to the subtitle generation unit, wherein the subtitle generation unit is configured to use the modification data as training data to train the subtitle generation model, wherein the trained subtitle generation model is configured to generate subtitle data for new content data by reflecting the modification data (Mowzoon: [0021], where the user can edit and upload the corrected text allowing the training model to retrain and reduce its errors with each such upload thereby become more and more accurate as time goes on; [0022], where the corrected text and accompanying signal is added to the training data pool allowing improvements and greater accuracy for Subsequent runs of the model).
Claim(s) 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Mowzoon in view of Parthasarathi and Skarbovsky, and further in view of Yegnanarayanan.
In regards to Claim 7, the combined teaching of Mowzoon, Parthasarathi, and Skarbovsky discloses the apparatus of claim 1, but fails to explicitly disclose wherein the subtitle modification unit is further configured to: manage the subtitle data in corpus units divided based on a predetermined criterion, extract a corpus corresponding to the modified portion of the modification data from the first subtitle data, and calculate the suitability to be higher when a similarity between an original of the extracted corpus and a modification to the corpus is higher.
Yegnanarayanan from a similar endeavor teaches managing the subtitle data in corpus units divided based on a predetermined criterion, extracting a corpus corresponding to the modified portion of the modification data from the first subtitle data, and calculating the suitability to be higher when a similarity between an original of the extracted corpus and a modification to the corpus is higher (Yegnanarayanan: [0159], where any suitable technique(s) may be used for determining which other portions of text are similar enough to the text portion whose extracted fact was corrected by the user to have a similar correction applied to that other portion of text; [0160], where any suitable technique(s) may be used to automatically apply a user's correction to extract similar facts from other text portions; [0162], where when the user-correction labeled text is added to the training data, it may be weighted differently from the original training data…the user correction training data may be weighted more heavily than the original training data, to ensure that it has an appreciable effect on the processing performed by the re-trained model. However, any suitable weighting may be used). Yegnanarayanan’s matching, searching, and comparing mechanism (see claims 4-5) operates over features of text, which correlates to corpus units. His similarity comparison between the original text and the user’s correction correlates to the similarity between an original of the extracted corpus and a modification to the corpus. Applying different weights to the original of the extracted corpus and the user-modified corpus correlates to calculating the suitability to be higher when a similarity between an original of the extracted corpus and a modification to the corpus is higher. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Mowzoon, Parthasarathi, and Skarbovsky in view of Yegnanarayanan such that their disclosed apparatus of claim 1 would also include Yegnanarayanan’s method of managing the subtitle data in corpus units divided based on a predetermined criterion, extracting a corpus corresponding to the modified portion of the modification data from the first subtitle data, and calculating the suitability to be higher when a similarity between an original of the extracted corpus and a modification to the corpus is higher. This would ensure that the user corrected text and original text are properly compared, so that the more similar they are, the more likely it is accurate and will be used for the subtitles.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Jin et al., US Pub. 2025/0157234 A1 teach using artificial intelligence to caption/describe images based on detected visual features.
Homyack et al., US Pub. 2015/0208139 A1 teach live captioning and caption extraction of an event, where the captioner can provide additional words during the process.
Sugimoto et al., US Pub. 2022/0027558 A1 teach extracting keywords from a text and generating weights for them based on the number of surrounding words associated with each keyword.
Snider et al., US Pub. 2015/0006199 A1 teach extracting facts from a medical text, wherein a user can make corrections to the extracted facts
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HANNA CHONG whose telephone number is (571)270-0520. The examiner can normally be reached Monday - Friday, 8 a.m. - 5 p.m. ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Nathan Flynn can be reached at (571) 272-1915. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HANNA CHONG/Examiner, Art Unit 2421
/NATHAN J FLYNN/Supervisory Patent Examiner, Art Unit 2421