DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 09/11/2026 has been entered.
Response to Arguments
Applicant's arguments filed 08/11/2026 have been fully considered but they are not persuasive.
Applicant argues Griffin does not disclose downloading, as language-dependent object, a translation file that is separate from the video file and that does not comprise video data of the video file.
Griffin et al teaches the audio portion of a piece of AV content may be automatically translated into multiple languages. Further, the audio portion of each individual speaker retains a sense of individuality when translated. As such, a user listening to any translated version will be able to distinguish the different speakers in the translated version of the content. In addition, the reference teaches separating the background noise audio data, the first speaker audio data, and the second speaker
audio data; translate the first speaker audio language to first speaker audio language of a second language; translate the second speaker audio language to second speaker
audio language of a second language; and generate encoded audio data including the first speaker audio language of the second language, the second speaker
audio language of the second language, and the background noise audio data; and a transmitter configured to transmit the encoded audio data to the content user device (Para. 0124-126; Claim 1). Wherein, the reference clearly teaches the audio portion of a piece of AV content and to transmit the encoded audio data to the content user device. The examiner relied on Olkha et al to meet limitations of downloading and the at least one/first translation file being stored separately from the video file, as discussed below.
Applicant argues Olkha fails to teach separating a reusable video file from a language-dependent translation file nor does Olkha teach that the claimed first translation file is downloaded from a target server corresponding to the first language type and synchronously played with the separately stored video file. A cloud source or translation service is not the claimed file- and-server architecture (Remarks: Page 13).
Olkha et al teaches a user device can stream audio from one cloud resource while downloading content from a second cloud resource and a user device can download content from multiple cloud resources for more efficient downloading. The media guidance application, implemented on control circuitry 404, requests a translation 110 to a language native to the user 112 from a remote cloud server 106. The translation 110 may comprise a subtitle track, closed-captioning track, and/or an audio recording in the native language. For example, the translation 110 may include an audio translation of the movie in English or Spanish. In another example, the translation 110 may include a subtitle file which includes an English or Spanish translation of the movie 104. The media guidance application, via control circuitry 404, may replace the audio file corresponding to the media asset 104 with a second audio file, where the second audio file is a translation of the audio file of the media asset 104 in a language native to the user 112 (Figure 5; Para. 0111, 0130-133) teaches a reusable video file from a language-dependent translation file and first translation file is downloaded from a target server corresponding to the first language type and synchronously played with the separately stored video file.
Applicant argues Ingel does not disclose a first or second translation file that is separate from the video file, nor does it disclose switching the language by continuing to play the same video file with a different translation file.
The examiner relied on Olkha et al to meet limitations of translation files are separate from the video files. Ingel et al teaches in response to a first received input, step 3010 may manipulate an aspect of a speech of the first person in the video, and in response to a second received input, step 3010 may manipulate an aspect of a speech of the second person in the video, where the second person may differ from the first person, and where the aspect of the speech of the first person may be the same as or different from the aspect of the speech of the second person. In some examples, at least one of the aspect of the speech of the first person and the aspect of the speech of the second person may comprise at least one of speech rhythm, speech tempo, pauses in speech, language, language register, and so forth (Para. 0438) meeting limitations switching the language by continuing to play the same video file with a different translation file.
In response to applicant’s argument that there is no teaching, suggestion, or motivation to combine the references, the examiner recognizes that obviousness may be established by combining or modifying the teachings of the prior art to produce the claimed invention where there is some teaching, suggestion, or motivation to do so found either in the references themselves or in the knowledge generally available to one of ordinary skill in the art. See In re Fine, 837 F.2d 1071, 5 USPQ2d 1596 (Fed. Cir. 1988), In re Jones, 958 F.2d 347, 21 USPQ2d 1941 (Fed. Cir. 1992), and KSR International Co. v. Teleflex, Inc., 550 U.S. 398, 82 USPQ2d 1385 (2007). In this case, Griffing et al teaches the audio portion of a piece of AV content and to transmit the encoded audio data to the content user device. Olkha et al teaches a user device can stream audio from one cloud resource while downloading content from a second cloud resource or a user device can download content from multiple cloud resources for more efficient downloading. Therefore, the examiner concluded that it would have been obvious to one of ordinary skill in the art to modify the reference before the effectively filing date of the claimed invention for the purpose of providing a translation of the media content in a language that is native to the user when the user is unable to sufficiently comprehend the media content in its original language (Para. 0002).
Similarly, Ingel et al teaches receiving a selection of a second language type by the first user; and switching from playing the video file and the first translation file to playing the video file and a second translation file corresponding to the second language type (Para. 0438). Thus, examiner concluded that it would have been obvious to one of ordinary skill in the art to modify the combination before the effectively filing date of the claimed invention for the purpose of providing a translation of the media content in a language that is native to the user when the user is unable to sufficiently comprehend the media content in its original language.
In response to applicant's argument that the references fail to show certain features of the invention, it is noted that the features upon which applicant relies (i.e., a downloaded translation file that is separate from the video file and excludes video data, nor the claimed synchronous playback and same-video-file switching workflow) are not recited in the rejected claim(s). Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993).
Olkha et al teaches concept of downloading and synchronizing (Para. 0048, 0111, claim 8). The reference also teaches the at least one/first translation file being stored separately from the video file (Figure 6; Para. 0111, 0130-132, Claim 8). The examiner relied on Ingel et al to teach same video file switching workflow (Para. 0438).
With respect to 9 and 24, the examiner took official notice in non-final action dated 12/29/2025. The applicant did not request any documentary evidence in response filed 03/30/2026. Applicant has failed to seasonably challenge the Examiner's assertions of well-known subject matter in the previous Office action(s) (Non-final 12/29/2025) pursuant to the requirements set forth under MPEP §2144.03. A "seasonable challenge" is an explicit demand for evidence set forth by Applicant in the next response (Remarks: 03/30/2026). Accordingly, the claim limitations the Examiner considered as "well known" in the first Office action (Non-final 12/29/2025) are now established as admitted prior art of record for the course of the prosecution. See In re Chevenard, 139 F.2d 71, 60 USPQ 239 (CCPA 1943).
To advance prosecution, the examiner has provided documentary evidence with respect to claims 9 and 24.
With respect to claims 6-7, 21-22, applicant argues Arsenault fails to teach a separately stored translation file that excludes video data and is synchronously plated with a reusable video file.
In response to applicant's argument that the references fail to show certain features of the invention, it is noted that the features upon which applicant relies (i.e., separately stored translation file that excludes video data) are not recited in the rejected claim(s). Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993).
In response to applicant's arguments against the references individually, one cannot show nonobviousness by attacking references individually where the rejections are based on combinations of references. See In re Keller, 642 F.2d 413, 208 USPQ 871 (CCPA 1981); In re Merck & Co., 800 F.2d 1091, 231 USPQ 375 (Fed. Cir. 1986). The examiner relied on Griffin, Olka and Ingel to meet limitations of a separately stored translation file that excludes video data and is synchronously plated with a reusable video file as discussed above.
In response to applicant’s argument that there is no teaching, suggestion, or motivation to combine the references, the examiner recognizes that obviousness may be established by combining or modifying the teachings of the prior art to produce the claimed invention where there is some teaching, suggestion, or motivation to do so found either in the references themselves or in the knowledge generally available to one of ordinary skill in the art. See In re Fine, 837 F.2d 1071, 5 USPQ2d 1596 (Fed. Cir. 1988), In re Jones, 958 F.2d 347, 21 USPQ2d 1941 (Fed. Cir. 1992), and KSR International Co. v. Teleflex, Inc., 550 U.S. 398, 82 USPQ2d 1385 (2007). In this case, Arsenault’s caching or distance related concept with motivation that it would have been obvious to one of ordinary skill in the art to modify the combination of Griffin, Olkha and Ingel before the effectively filing date of the claimed invention for the purpose of providing content that is not susceptible to a buffer stall, delay and degradation of quality meeting claimed limitations.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 8, 14, 17-20 and 23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Griffin et al (US PG Pub No. 2024/0211704), in view of Olkha et al (US PG Pub No. 2020/0007946), further in view of Ingel et al (US PG Pub No. 2021/0224319).
Regarding claims 1 and 14, Griffin et al teaches a method of cross-language video processing (Abstract), comprising:
in response to a video playing request by a first user (S208) (i.e. receiving request for dubbed content) (Figure 9; Para. 0117), obtaining a first video to be played (Para. 0119), the first video being a video posted by a second user (i.e. content provider provides video(s) to service providing systems and audio processing systems) (S204) (Figure 1A; Para. 0020, 0034);
obtaining a video file associated with the first video, and at least one translation file (i.e. Spanish-dubbed version of the video corresponding to the video snapshot 902) (Abstract, Para. 0122, 0125-126, Claim 1), the at least one translation file being obtained based on an original audio file (i.e. audio data) of the first video being translated according to a respective language type of at least one language type (Figure 4; Para. 0020, 0075), and the original audio file and the video file being obtained by decoupling the first video (i.e. speech processor 316 to separate audio data from the original AV content received from content provider) (Figure 4; Para. 0075);
obtaining first language type matching play demand information of the first user, according to the respective language type corresponding to the at least one translation file (i.e. AV device 110 transmits a request 122 to service providing system 106 to provide the version of the video corresponding to video snapshot 902 in Spanish) (Figure 8-9; Para. 0122, 0125-126, Claim 1); and
a first translation file from at least one target server corresponding to the first language type, and playing the video file and the first translation file (Para. 0124-126, claim 1).
The reference teaches AV device comprising memory and is able to play dubbed content. However, fails to explicitly state downloading and synchronizing and the at least one/first translation file being stored separately from the video file.
In similar field of endeavor, Olkha et al teaches concept of downloading and synchronizing (Para. 0048, 0111, claim 8). The reference also teaches the at least one/first translation file being stored separately from the video file (Figure 6; Para. 0111, 0130-132, Claim 8). Therefore, it would have been obvious to one of ordinary skill in the art to modify the reference before the effectively filing date of the claimed invention for the purpose of providing a translation of the media content in a language that is native to the user when the user is unable to sufficiently comprehend the media content in its original language (Para. 0002).
The combination teaches during playing the video file and the first translation file (See Griffin: Para. 0027, 0122). The combination is unclear with respect to receiving a selection of a second language type by the first user; and switching from playing the video file and the first translation file to playing the video file and a second translation file corresponding to the second language type.
In similar field of endeavor, Ingel et al teaches receiving a selection of a second language type by the first user; and switching from playing the video file and the first translation file to playing the video file and a second translation file corresponding to the second language type (Para. 0438). Therefore, it would have been obvious to one of ordinary skill in the art to modify the combination before the effectively filing date of the claimed invention for the purpose of providing a translation of the media content in a language that is native to the user when the user is unable to sufficiently comprehend the media content in its original language.
Regarding claims 2 and 17, Griffin, Olkha and Ingel, the combination teaches wherein obtaining the first language type matching the play demand information of the first user, according to the respective language type corresponding to the at least one translation file comprises:
displaying the at least one language type on a playing interface of the first video (Griffin: Figure 9); and in response to selecting by the first user on the at least one language type (Griffin: Para. 0119), obtaining the first language type selected by the first user (Griffin: Para. 0121-122).
Regarding claims 3 and 18, Griffin, Olkha and Ingel, the combination teaches wherein obtaining the first language type matching the play demand information of the first user, according to the respective language type corresponding to the at least one translation file comprises:
obtaining an application language of the first user according to user information of the first user; and obtaining the first language type as a language type of the application language, according to the respective language type corresponding to the at least one translation file (i.e. determines, based on the profile, whether the language is a language native to the user) (Olkha: Para. 0116-117).
Regarding claims 4 and 19, Griffin, Olkha and Ingel, the combination teaches wherein the translation file comprises a track audio file, and obtaining the video file associated with the first video, and the at least one translation file comprises:
in accordance with a determination that the first video has multi-track playing permission, obtaining the video file associated with the first video, and at least one track audio file (i.e. application that allows a user to select a media asset, compares a language proficiency level required to comprehend the media asset and the language proficiency level of the user to determine whether to automatically provide translations, and selectively delivers an output comprising the translation) (Olkha: Para. 0020, 0048, 0130).
Regarding claims 5 and 20, Griffin, Olkha and Ingel, the combination teaches wherein displaying a multi-track switch on a playing interface of the first video (Griffin: Fig. 9); and in response to triggering by the first user on the multi-track switch (Griffin: Figure 9; Para. 0119-120), determining that the first video has multi-track play permission, the multi-track play permission being used to enable playing permission of a track audio file of the first video, corresponding to the respective language type of the at least one language type (i.e. application that allows a user to select a media asset, compares a language proficiency level required to comprehend the media asset and the language proficiency level of the user to determine whether to automatically provide translations, and selectively delivers an output comprising the translation) (Olkha: Para. 0020, 0048, 0130).
Regarding claims 8 and 23, Griffin, Olkha and Ingel, the combination teaches wherein downloading the target translation file from the target server corresponding to the target language type comprises:
obtaining, according to the target server, a target download address of the first video, corresponding to the first language type; and downloading the first translation file using the target download address (Griffin: Figure 8; Para. 0114-115 and Olkha: Para. 0048, 0111, claim 8).
Claim(s) 6-7 and 21-22 is/are rejected under 35 U.S.C. 103 as being unpatentable over Griffin et al, in view of Olkha et al, in view of Ingel et al, further in view of Arsenault et al (US PG Pub No. 2018/0241836).
Regarding claims 6 and 21, Griffin, Olkha and Ingel, the combination teaches wherein before downloading the first translation file matching the first language type, from the target server corresponding to the first language type: in accordance with a determination, by querying, that the first target translation file corresponding to the first language type as discussed above. The combination is unclear with respect stored in a local server, determining that the local server is the target server.
In similar field of endeavor, Arsenault et al teaches stored in a local server, determining that the local server is the target server (i.e. Client device plays chunks for one or more segments from buffer and/or edge server) (Para. 0041-42). Therefore, it would have been obvious to one of ordinary skill in the art to modify the combination before the effectively filing date of the claimed invention for the purpose of providing content that is not susceptible to a buffer stall, delay and degradation of quality.
Regarding claims 7 and 22, Griffin, Olkha and Ingel, the combination teaches wherein before downloading the first translation file from the target server corresponding to the first language type; querying at least one server associated with the first language type of the first video, as discussed above.
The combination is unclear with respect to obtaining, from the at least one server, a target server with a minimum transmission distance from the first user.
In similar field of endeavor, Arsenault et al teaches determining, from the at least one server, a target server with a minimum transmission distance from the first user (i.e. Client device plays chunks for one or more segments from buffer and/or edge server) (Figure 2; Para. 0041-42). Therefore, it would have been obvious to one of ordinary skill in the art to modify the combination before the effectively filing date of the claimed invention for the purpose of providing content that is not susceptible to a buffer stall, delay and degradation of quality.
Claim(s) 9 and 24 is/are rejected under 35 U.S.C. 103 as being unpatentable over Griffin et al, in view of Olkha et al, in view of Ingel et al, further in view of Ojiro et al (US PG Pub No. 2019/0387271).
Regarding claims 9 and 24, Griffin, Olkha and Ingel, the combination teaches wherein downloading the first translation file and synchronously playing the video file and the first translation file (Griffin: Para. 0122 and Olkha: Para. 0048, 0111, claim 8) comprises:
obtaining at least one translation file segment of the first translation file matching the first language type and a segment order corresponding to a respective translation file segment of the at least one translation file segment (Griffin: Figures 6-7; Para. 0108-109);
downloading, the translation file segment based on the segment order of the respective translation file segment (Griffin: Para. 0111);
coupling the downloaded translation file segment to a video file segment corresponding to the video file to obtain a target video segment, the downloaded translation file and the video file segment having a same timestamp (Griffin: Figure 6-7; Para. 0076, 0100, 0110-111 and Olkha: Para. 0048, 0111, claim 8); and
playing the target video segment to synchronously play the translation file segment in the first translation file and the video file segment in the video file (Griffin: Griffin: Para. 0122 and Olkha: Para. 0048).
The combination is unclear with respect to sequentially downloading and precaching.
In similar field of endeavor, Ojiro et al teaches to sequentially downloading and precaching (Figure 1, 11; Abstract, Para. 0044, 0183, 0197-199). Therefore, it would have been obvious one of ordinary skill in the art is to modify the combination by specifically providing sequentially downloading and precaching before the effectively filing date of the claimed invention for the common knowledge purpose of providing user content accurately and in timely manner without interruption.
Claim(s) 10-11 is/are rejected under 35 U.S.C. 103 as being unpatentable over Griffin et al (US PG Pub No. 2024/0211704), in view of Olkha et al (US PG Pub No. 2020/0007946).
Regarding claim 10, Griffin et al teaches a method of cross-language video processing, comprising:
in response to a video posting request by a second user, obtaining a first video to be posted (i.e. content provider provides video(s) to service providing systems and audio processing systems) (S204) (Griffin: Figure 1A; Para. 0020, 0034);
decoupling the first video to be posted, to obtain an original audio file and a video file (i.e. audio portion of the AV content are separated) (Griffin: Para. 0020); converting the original audio file into a translation file based on a respective language type of at least one language type to obtain at least one translation file (text of each language is converted to speech) (Griffin: Para. 0020);
posting the video file, the original audio file and the at least one translation file, to store the at least one translation file in at least one server matching the respective language type corresponding to the at least one translation file (i.e. Memory of the audio processing system stores respective language type corresponding to translations) (Griffin: Figure 8; Para. 0020, 0034, 0114-115); and
wherein the video file and a first translation file of the at least one translation file are used for play in response to a cross-language video processing request by a first user, and the first translation file matches a play demand of the first user (Griffin: Figure 1A; Para. 0020, 0034).
The reference teaches AV device comprising memory and able to play dubbed content. However, fails to explicitly state synchronizing and at least one file separately from the video file.
In similar field of endeavor, Olkha et al teaches concept synchronizing (Para. 0048, 0111, claim 8) and at least one file separately from the video file (Figure 6; Para. 0111, 0130-132, Claim 8). In addition, the reference also teaches posting the video file, the original audio file and the at least one translation file, to store the at least one translation file in at least one server matching the respective language type corresponding to the at least one translation file (Para. 0048, 0111, claim 8). Therefore, it would have been obvious to one of ordinary skill in the art to modify the reference before the effectively filing date of the claimed invention for the purpose of providing a translation of the media content in a language that is native to the user when the user is unable to sufficiently comprehend the media content in its original language (Para. 0002).
Regarding claim 11, Griffin and Olkha, the combination teaches wherein storing the at least one translation file in the at least one server matching the respective language type corresponding to the at least one translation file (Griffin: Para. 0111-112 and Olkha: Para. 0048) comprises:
obtaining a language application region corresponding to a respective translation file of the at least one translation file, according to the language type corresponding to the respective translation file (Olkha: Para. 0049, 0062, 0116-117);
querying at least one server associated with the language application region corresponding to the respective translation file (Griffin: Para. 0111-112 and Olkha: Para. 0062-63); and
sending the respective translation file to the server associated with the corresponding language application region (Griffin: Para. 0111-112 and Olkha: Para. 0062-63).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KUNAL LANGHNOJA whose telephone number is (571)270-3583. The examiner can normally be reached M-F: 9:00AM - 5:00PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Brian Pendleton can be reached at (571) 272-7527. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KUNAL LANGHNOJA/Primary Examiner, Art Unit 2425