DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-9 and 10-20 are rejected under 35 U.S.C. 102(a)(2) as being described by Li et al (US 2024/0064367).
For claim 1, Li et al teach a method for content generation, comprising:
in response to an audio edit request, presenting an audio edit panel comprising at least an audio generating control (e.g. paragraph 62: displaying an audio panel in the video editing interface. figure 1, paragraphs 35: In step 101, in response to a triggering operation acting on a first control in a video editing interface, a first panel is displayed. The first panel includes a first operation region for a first object. Paragraph 36: The first object may be text, a voice, an animation, etc. The user interacts with the system in the first operation region to operate the first object. For example, the user may input text or a voice through the first operation region, or select existing text or existing voice through the first operation region.);
in response to detecting a trigger on the audio generating control, obtaining a first text and a first audio corresponding to the first text (e.g. paragraph 38: the first control is a text-to-speech control, the first object is text, and the first operation region is a text input region; accordingly, in an embodiment, step 101 may include: in response to a triggering operation acting on the text-to-speech control in the video editing interface, displaying a first panel, where the first panel includes a text input region.), at least one of the first text or the first audio being determined based on a content entity to be edited (e.g. paragraph 46: In a case of text-to-speech, the resultant voice is added to the target video and synthesized with the target video, so that the voice corresponding to the inputted text content may be presented during a process of playing the target video.); and
adding the first text and the first audio into the content entity to obtain a first video, wherein the first text is presented overlapped on the content entity, and the first audio is configured to be at least a part of an audio corresponding to the first video (e.g. paragraph 58: As shown in FIG. 3, which illustrates a schematic diagram of a video editing interface, the text content 510 is the text content that has been added to the target video. The user selects the text content 510 and then triggers the first control 520. In an embodiment, the first control 520 may be a text-to-speech control…Each time the user switch the selected timbre, the system re-reads the text content with the selected timbre and may play the video clip (to which the target audio is to be added) in the target video that matches the position for adding the target audio in the preview window 630, to enable the user to preview the effect of adding the target audio to the target video).
Claims 11 and 20 are rejected for the same reasons as discussed in claim 1 above, wherein paragraph 89 of Li et al teach processor execute programs stored in memory.
For claims 2 and 12, Li et al teach at least one of the first text or the first audio is generated based on the content entity using a machine learning model (e.g. paragraph 66: inputting the phoneme and rhythm symbol information into a deep learning model to obtain multiple voice clips corresponding to the first object with the first feature).
For claims 4 and 14, Li et al teach at least one of a text style of the first text or the first timbre type of the first audio is determined based on the content entity using the machine learning model (e.g. paragraph 66: rhythm-and-texture fusion information into phoneme and rhythm symbol information; inputting the phoneme and rhythm symbol information into a deep learning model to obtain multiple voice clips corresponding to the first object with the first feature. The determining the multiple beat positions of the background audio in the target video includes: acquiring the multiple beat positions of the background audio in the target video by using a beat detection model).
For claims 3 and 13, Li et al teach the first audio is generated by performing text-to-speech on the first text based on a first timbre type (e.g. paragraph 65: in a case of text-to-speech, adding a second object with a second feature to the target video (i.e., adding a target audio with a target timbre to the target video, where the target audio is generated for the text content) ), and wherein the first timbre type is determined by at least one of the following: determining a timbre type specified by a user as the first timbre type, or selecting the first timbre type from a timbre library randomly (e.g. paragraph 54: In response to the user selecting the purity timbre…target video is played in the preview window, while playing the audio of “clouds are floating, and the deer is running happily” in the purity timbre simultaneously.).
For claims 5 and 15, Li et al teach presenting an adjustment control in association with the first video, the adjustment control indicating adjustment to the first text and the first audio in the content entity; in response to detecting a trigger on the adjustment control, obtaining a second text and a second audio for replacing the first text and the first audio, respectively, at least one of the second text or the second audio being determined based on the content entity; and adding the second text and the second audio into the content entity to obtain a second video, wherein the second text is presented overlapped on the content entity, and the second audio is configured to be at least a part of an audio corresponding to the second video (e.g. paragraph 62: When the user selects a specific timbre, the text content is read with the currently selected timbre. When the user confirms, the timbre of the existing target audio for the text content is changed to the timbre currently selected by the user. In other words, with the audio panel).
For claims 6 and 16, Li et al teach obtaining the second text and the second audio comprises: in response to detecting a trigger on the adjustment control, presenting at least one of a text style selection entry or a timbre selection entry; receiving at least one of a selection of a second text style via the text style selection entry, or a selection of a second timbre type via the timbre selection entry; and obtaining the second text with the second text style and the second audio with the second timbre type for replacing the first text and the first audio, respectively (e.g. paragraph 62: When the user selects a specific timbre, the text content is read with the currently selected timbre. When the user confirms, the timbre of the existing target audio for the text content is changed to the timbre currently selected by the user. In other words, with the audio panel. Figure 3 shows “Text” icon is underlined and “style” icon next to the “Text-to-speech” icon).
For claims 7 and 17, Li et al teach in response to a trigger on an edit operation of the first text, presenting a text edit box corresponding to the first text; receiving an updated third text via the text edit box; presenting the third text overlapped on the content entity Figure 3 shows “Text” icon is underlined and “style” icon next to the “Text-to-speech” icon); and removing the first audio from the content entity (e.g. figure 3: “Text” icon is underlined and “Delete” icon is right next to the “style’ icon).
For claims 8 and 18, Li et al teach the acts further comprise: in response to detecting a text-to-speech request for the third text, obtaining a third audio corresponding to the third text; and adding the third audio into the content entity to obtain a third video (e.g. figure 3, “Text” icon is underlined and “text-to-speech” 520 icon is above the “Text” icon).
For claims 9 and 19, Li et al teach in one or more times of presentation, a visual style of the audio generating control is randomly selected from a plurality of candidate visual styles (e.g. figure 4 shows plurality of timbre styles can be selected by user).
For claim 10, Li et al teach obtaining the first text comprises: sampling a part of content from the content entity; extracting, based on the part of content, first semantic information corresponding to the part of content using a semantic model; and generating the first text based on the first semantic information and prompt information, using a machine learning model.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Li et al, as applied to claims 1-9 and 11-20 above and further in view of Zhang et al (US 2022/0415071).
For claim 10, Li et al do not teach sampling a part of content from the content entity; extracting, based on the part of content, first semantic information corresponding to the part of content using a semantic model; and generating the first text based on the first semantic information and prompt information, using a machine learning model. Zhang et al teach sampling a part of content from the content entity; extracting, based on the part of content, first semantic information corresponding to the part of content using a semantic model; and generating the first text based on the first semantic information and prompt information, using a machine learning model (e.g. paragraph 53: feature-extraction is performed on the sample text to obtain semantic features of the sample text, and the basic network is trained based on the semantic features, so as to enable the basic network to learn the ability of extracting text content based on the semantic features, thereby obtaining the text recognition model). It would have been obvious to one ordinary skill in art to incorporate the teaching of Zhang et al into the teaching of Li et al to improve the accuracy and reliability of the text recognition performed by the text recognition model (e.g. paragraph 72, Zhang et al).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Wu et al (US 2026/0011346, see figure 3, step 330 obtain text input, step 340 generate video based on the input text) .
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DAQUAN ZHAO whose telephone number is (571)270-1119. The examiner can normally be reached M-Thur: 7:00 am-5:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Thai Tran can be reached on 571-272-7382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
Email: daquan.zhao1@uspto.gov.
Phone: (571)270-1119
/DAQUAN ZHAO/Primary Examiner, Art Unit 2484