DETAILED ACTION
This action is responsive to amendments filed on August 10th, 2026.
Claims 1~8 and 20~28 are examined.
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments with respect to claims 1~8 and 20~28 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claims 1~7 and 20~27 are rejected under 35 U.S.C. 103 as being unpatentable over Coudurier et al. hereinafter Coudurier (U.S 2013/0322854), Padi et al. hereinafter Padi (U.S 2011/0231180), Wilson et al. hereinafter Wilson (U.S 10,321,174) and Stankiewicz et al. hereinafter Stankiewicz (U.S 2016/0322080) in view of Stone (U.S 2023/0300399).
Regarding Claim 1,
Coudurier taught a method, comprising:
transmitting the one or more sets of caption information to a first client device via a first route [¶27, captions delivered out-of-band]; and
transmitting the video content associated with the one or more sets of caption information to a second client device via a second route different than the first route [¶29, video file 110 is sent to user device 104 in-band].
Coudurier did not specifically teach receiving, from at least one external source, one or more sets of caption information in the predicted language of the user, wherein the one or more sets of caption information is separate from closed-caption data embedded within or linked to the video content.
Padi taught in step 325, media device 110 transmits a request for closed captioning for the media content selected to IPG server 130, which then determines whether closed captioning is available for the media content in the requested language [¶29]. Step 345 may follow step 325 when it is determined that the language for closed captioning requested by a user is not available via IPG server 130 [¶33]. Following step 345, in step 350, media device 110 obtains from translation service 140 in translation server 145, a translation for a specified portion of media content being displayed to a user [¶34].
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention was made, to combine, Padi’s teaching with the teachings of Coudurier because the combination would improve the availability of additional captions in a language of interest to the user [Padi: ¶1].
The combination of Coudurier and Padi did not specifically teach determining, by a server, a predicted language to present closed caption data associated with video content to a user based on at least one of a subscription history of the user, a user profile, or a viewing history of the user.
Wilson taught determining, by a server, a predicted language to present closed caption data associated with video content to a user based on at least one of a subscription history of the user, a user profile, or a viewing history of the user (C8: 58~66, the account preferences 110 are established by default based on the account profile data 108 for each account or user. For example, the server computer 104 may have access to metadata that specifies the most popular languages for each country. The server computer 104 then approximates, based on the user's country of origin specified in the account profile data 108 and the metadata which languages the user is likely to be familiar with and populates the defaults for the account preferences 110).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention was made, to combine, Wilson’s teaching with the teachings of Coudurier and Padi because the combination would conveniently or efficiently address user preferences for subtitles and dubbing (Wilson: C2: 16~18).
The combination of Coudurier, Padi, and Wilson did not specifically teach wherein the video content received by the second client device via the second route is transferred from the second client device to the first client device over a local connection that is separate from the first route and the second route.
Stankiewicz taught wherein the video content received by the second client device via the second route is transferred from the second client device to the first client device over a local connection that is separate from the first route and the second route [¶13, device 102 is communicatively coupled to a display device 104; ¶15, media source 108 may provide any one or more of video content 112(1), which has associated out-of-band closed caption data. Out-of-band closed caption data is delivered in a separate file distinct from a file containing the video content thus, showing that video].
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention was made, to combine, Stankiewicz’s teaching with the teachings of Coudurier, Padi, and Wilson, because the combination would allow a single text renderer to provide a consistent look and feel to closed caption data originating in any of multiple formats (abstract).
The combination of Coudurier, Padi, Wilson, and Stankiewicz did not specifically teach wherein the first client device synchronizes the video content received over the local connection with the one or more sets of caption information received via the first route by: identifying one or more images of the video content associated with mouth movements of a speaker shown in the video content, analyzing the one or more images by applying a trained model having a plurality of parameters with different weighting values, the plurality of parameters including a first parameter corresponding to the mouth movements of the speaker and having a first weighting value and a second parameter corresponding to a size of a mouth of the speaker and having a second weighting value, and adjusting a display timing of the one or more sets of caption information relative to the video content such that the one or more sets of caption information and the video content are synchronized.
Stone taught wherein the first client device synchronizes the video content received over the local connection with the one or more sets of caption information received via the first route by: identifying one or more images of the video content associated with mouth movements of a speaker shown in the video content, analyzing the one or more images by applying a trained model having a plurality of parameters with different weighting values, the plurality of parameters including a first parameter corresponding to the mouth movements of the speaker and having a first weighting value and a second parameter corresponding to a size of a mouth of the speaker and having a second weighting value, and adjusting a display timing of the one or more sets of caption information relative to the video content such that the one or more sets of caption information and the video content are synchronized [¶19, AI or machine learning may be used (e.g., by encoder/packager 112 or user devices 116) to align and sync audio 106 and/or video 108 with closed captions 110. For example, encoder/packager 112 or user devices 116 may implement a software algorithm that listens to audio 106 and/or processes video 108 to determine when words being spoken (“mouth movements”) in content 104 match those of closed captions 110]. Note: A person of ordinary skill in the art would have found it obvious to apply AI/ML/Neural Network models on mouth size to align and sync audio and/or video with captions, as both the movement and size of a mouth are indicative of words being spoken thus, the results would have been predictable. Further, machine learning algorithms are also well-known to involve parameters associated with mouth movement or shape along with mouth size and associated weighted values for the purposes of training the model to predict the audio/speech timeline and align the caption to a video content).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention was made, to combine, Stone’s teaching with the teachings of Coudurier, Padi, Wilson, and Stankiewicz, because the combination would improve viewing experience [Stone: ¶1].
Regarding Claim 2,
Coudurier taught further comprising combining the video content and the one or more sets of caption information by the first client device [¶24, user device 114 receives caption file 108 and video file 110 and renders them on media player 116].
Regarding Claim 3,
Coudurier-Padi-Wilson-Stankiewicz taught further comprising combining the video content and the one or more sets of caption information by the second client device [Stankiewicz: ¶13, device 102, which includes, or is communicatively coupled to, a display device 104; ¶17, text renderer 130 renders the closed caption text 132 for display via display device 104]. The rationale to combine as discussed in claim 1, applies here as well.
Regarding Claim 4,
Coudurier taught further comprising transmitting the combined video to a displaying device for viewing [¶15, user device 114 may be a television].
Regarding Claim 5,
Coudurier-Padi-Wilson-Stankiewicz taught wherein the first route includes communications via an Internet device [Stone: ¶14, network 110 can include a cable television network, radio frequency (RF), microwave, satellite, and/or data network, such as the Internet]. The rationale to combine as discussed in claim 1, applies here as well.
Regarding Claim 6,
Coudurier-Padi-Wilson-Stankiewicz taught wherein the second route includes communications via a satellite [Stone: ¶14, network 110 can include a cable television network, radio frequency (RF), microwave, satellite, and/or data network, such as the Internet]. The rationale to combine as discussed in claim 1, applies here as well.
Regarding Claim 7,
Coudurier taught wherein the one or more sets of caption information include a customized caption information based on the user information, and wherein the customized caption information is generated in response to an event that the server identifies an absence of closed captioned data in a specific language [¶27, if additional caption files 108 need to be added, such as for a different language, for the media program, a separate caption file 108 may be generated for a new language without affecting video file 110].
Regarding Claims 20~27, the claims are similar in scope to claims 1~7 and therefore, rejected under the same rationale.
Claims 8 and 28 are rejected under 35 U.S.C. 103 as being unpatentable over Coudurier, Padi, Wilson, Stankiewicz, and Stone in view of Reitan (U.S 2012/0316860).
Regarding Claims 8 and 28,
Coudurier-Padi-Wilson-Stankiewicz-Stone-Reitan taught wherein the customized caption information is generated by translating from an existing caption in an existing language [Reitan: ¶11, the caption translation system provides translated captions from a source caption into the target language].
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention was made, to combine, Reitan’s teaching with the teachings of Coudurier, Padi, Wilson, Stankiewicz, and Stone, because the combination would provide digital video to a much broader audience where language is no longer a barrier for anyone viewing the video after captions have been translated.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HEE SOO KIM whose telephone number is (571)270-3229. The examiner can normally be reached M-F 9AM-5PM.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Nicholas Taylor can be reached on (571) 272-3889. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HEE SOO KIM/Primary Examiner, Art Unit 2443