DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Preliminary Amendment
According to the preliminary amendment filed 09/18/2025, claims 1-11, 13-21 are pending, and claims 12, 22-23 have been canceled.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 10-11, 13, 15, 17-21 are rejected under 35 U.S.C. 103 as being unpatentable over Ellis et al. (US 20150312640) in view of Longan et al. (US 20190206128).
Note: all documents that are directly or indirectly incorporated by references in their entireties in Ellis (see paragraphs 0001, 0119, 0124, 0132, 0135, 0141, 0144, 0147, 0150, 0152, 0156, 0159, 0179, 0189, 0192, 0215, 0217, 0224, 0233) including Ser. No. 09/356,270 (corresponding to US 20050262542 -referred to as DeWeese), or in Logan (see paragraphs 0102, 0105, 0107, 0125, 0133) including 20100153885 (Yates) are respectively treated as part of the specification of Ellis or Longan (see for example, MPEP 2163.07 b).
Regarding claim 1, Ellis discloses a method of generating real-time video content (content with real time data from real time data source – see figures 1a, 34-35, 53D, DeWeese: figures 9, 15A, 16) , performed by at least one processor (processor at server, distribution facility or television receiving equipment – see include, but are not limited to, Ellis: paragraphs 0101, 0126, 0132) , the method comprising:
generating text content based on a broadcast outline (generating a text description/closed caption of content based on a broadcast information of news, program, etc. – see include, but are not limited to, Ellis: figures 1A, 3, 31, 34, paragraphs 0095, 0098, 0129, 0130, 0183; DeWeese: figures 15A-16);
generating a first voice (generating audio/sound/voice – see include, but are not limited to, Ellis: paragraphs 0092, 0095, 0233; DeWeese: paragraphs 0055, 0091, 0101);
generating a first view based on at least one of the text content or the first voice (generating a first view/display based on at least one of the text description/closed caption or the first audio of program/content– see include, but are not limited to, see Ellis: figures 1a, 34-35, 53C-53D, paragraphs 0095, 0098, 0221; DeWeese: figures 15a-16, paragraphs 0091, 0101-0103) ; and
transmitting a first video comprising the first voice and the first view in a real-time streaming manner (transmitting a first video comprising a first voice/audio and the first view in real time manner such as during broadcasting/viewing of the content – see include, but are not limited to, Ellis: figures 1, 34-35, 53C-53D, paragraphs 0095, 0098, 0201; DeWeese: figures 15A-16, paragraphs 0101-0103, 0105-0107) .
Ellis does not explicitly disclose generating a first voice based on the text content.
Logan discloses generating text content of a broadcast outline (generating textual information /closed captioning of events in virtual reality media assets/content – see paragraphs 0031, 0187, 0222);
generating a first voice based on the text content (generating audio of a second voice (e.g., of Martin Tyler) reading the textual information retrieved from closed-captioning – see include, but are not limited to, figure 13, paragraphs 0031, 0054, 0187; Yates: paragraphs 0053, 0126);
In addition to Ellis, Logan discloses generating first view based on the text or first voice and transmitting a first video comprising a first voice and the first view (generating a first view of object based on the text or selected voice and transmitting a first video comprising the selected voice and the first object – see include, but are not limited to, figures 1B, 7, 13, paragraphs 0054, 0187).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Ellis with the teaching of generating a first voice based on text content as taught by Longan in order to yield predictable result of allowing user to hear desired voice of the user (see paragraphs 0030-0031).
See also Zito, Jr. (US 20160021412: e.g., paragraphs 0083, 0086, 0101-0102) for teaching of providing different views with difference voices to user(s) to attract users.
Regarding claim 10, Ellis in view of Logan discloses the method as claimed in claim 1, wherein the generating of the text content based on the broadcast outline comprises: generating the text content based on a broadcast outline in a text format using a story generation model (story generation model is read on model/software that is used to generate a text format with titles, description, playlist, etc. on a display based on genre, time, program content, etc. – see include, but are not limited to, Ellis: figures 2-3, 31-32, paragraphs 0183, 0231, 0241; Longan: figures 3-4; Yates: figures 4, 6) .
Regarding claim 11, Ellis in view of Logan discloses the method as claimed in claim 1, wherein the broadcast outline is a song list comprising a plurality of song titles (song list of album title – see for example, Ellis: figures 1a, 54C-54D), and
the generating of the text content based on the broadcast outline comprises: normalizing the plurality of song titles (receiving/normalizing the plurality of song titles of a niche hub, album, etc.– see include, but are not limited to, Ellis: figures 1a, 54c, 54d);
collecting information associated with a plurality of songs included in the song list based on the normalized plurality of song titles using a search engine (collecting information such title, artist, album information, etc. based on information of the plurality of song title/tracks using a search engine - see include, but are not limited to, Ellis: figures 1a, 3, 53E-54D, paragraphs 0219, 0222, 0224; Yates: figure 1);
summarizing the collected information using a summarization model (summary the collected information of searched title/album/artist using a model listing songs/titles for the searched album, artist, type of music, etc. - see include, but are not limited to, Ellis: figures 1a, 3, 53E-54D, paragraphs 0219, 0222, 0224; Yates: figure 1); and
changing a style of the summarized information using a style conversion model (changing a style of summarized information based on type of music, name of album, artist, etc. using conversion model to display titles/songs based on type of music, album, artist, etc. - - see include, but are not limited to, Ellis: figures 1a, 3, 53E-54D, paragraphs 0219, 0222, 0224; Yates: figures 3-6).
Regarding claim 13, Ellis in view of Logan discloses the method as claimed in claim 1, wherein the broadcast outline is generated based on a collection result of collecting popular news information using a web that provides news (see include, but are not limited to, Ellis: figures 22, 32, 34, paragraphs 0095, 0185, 0251).
Regarding claim 15, Ellis in view of Logan discloses the method as claimed in claim 1, wherein the generating of the text content based on the broadcast outline comprises generating, using a voice recognition model, the text content based on a broadcast outline in a sound format (software or model that recognizes or detects a text content such as closed captioning or other text in sound/audio portion – see include, but are not limited to, Longan: paragraphs 0031, 0054, 0222).
Regarding claim 17, Ellis in view of Logan discloses the method as claimed in claim 1, further comprising:
acquiring a viewer's real-time chat for the first video transmitted in the real-time streaming manner (see include, but are not limited to, Ellis: figures 50-51, paragraphs 0106, 0216, 0217; DeWeese: figures 15A-16);
modifying the broadcast outline based on at least a part of the real-time chat (modifying/adding real time chat to program/channel - see include, but are not limited to, Ellis: figures 50-51, paragraphs 0106, 0216, 0217; DeWeese: figures 15A-16);
generating modified text content based on the modified broadcast outline using a story generation model (see include, but are not limited to, Ellis: figures 50-51, paragraphs 0106, 0216, 0217; DeWeese: figures 15A-16 – story generation model is read on software or model for modifying/changing text content with chat or message from user - see include, but are not limited to, Ellis: figures 50-51, paragraphs 0106, 0216, 0217; DeWeese: figures 15A-16; Logan: paragraphs 0031, 0054, 0187);
generating a third voice and a third view based on the modified text content (generating a voice of user/or different user and a view with voice/sound/audio associated with another user reaction to the time of the event of program – see include, but are not limited to, Logan: figures 1b, 15-16, paragraphs 0031, 0054, 0187; DeWeese: figures 15a-16); and
transmitting a third video comprising the third voice and the third view in a real-time streaming manner (transmitting another video comprising a different voice/audio/sound of the user with the different view in real time manner for displaying – see include, but are not limited to, Logan: figures 1b, 15-16, paragraphs 0031, 0054, 0187; DeWeese: figures 15a-16).
Regarding claim 18, Ellis in view of Logan discloses the method as claimed in claim 1, wherein the first view comprises at least one of a virtual streamer in a character form (virtual object/avatar 124 or 122 in a character form – see Logan: figure 1B, 10, paragraphs 0025, 0053-0055, 0087, 0089) a virtual streamer in a virtual person form (avatar of user/friend), or a background view associated with the text content see Logan: figure 1B, 10, paragraphs 0025, 0053-0055, 0087, 0089).
Regarding claim 19, Ellis in view of Logan discloses the method as claimed in claim 1, wherein the first view comprises a virtual streamer in a character form or a virtual person form (first view comprises object avatar 122/124 - see Logan: figure 1B, 10, paragraphs 0025, 0053-0055, 0087, 0089), and the generating of the first view based on at least one of the text content or the first voice comprises generating the first view comprising a facial expression change (facial expression) of the virtual streamer reflecting an emotion associated with the text content, using an emotion prediction model and a facial expression change model (see include, but are not limited to, Logan: figure 1B, paragraphs 0014-0018, 0153-0156, 0185, 0225).
Regarding claim 20, Ellis in view of Logan discloses the method as claimed in claim 1, wherein the first view comprises a virtual streamer in a character form or a virtual person form, and
the generating of the first view based on at least one of the text content or the first voice comprises generating the first view comprising a mouth shape change of the virtual streamer (widening of mouth classified as being shock) , who is speaking the first voice, using a talking head model (widening mouth, moving head, etc. of the user who is talking/speaking - see include, but are not limited to, Logan: paragraphs 0014-16, 0026, 0153-0155, 0178).
Regarding claim 21, Ellis in view of Logan discloses the method as claimed in claim 1, wherein the first view comprises a virtual streamer in a character form or a virtual person form, and the generating of the first view based on at least one of the text content or the first voice comprises generating the first view comprising a gesture of the virtual streamer using a gesture model (- see include, but are not limited to, Logan: paragraphs 0014-16, 0026, 0153-0155, 0178 wherein “gesture” is interpreted as any gesture or motion of mouth, eyes, head, hand, or other part of body of user).
Claims 2-7 are rejected under 35 U.S.C. 103 as being unpatentable over Ellis et al. (US 20150312640) in view of Longan et al. (US 20190206128) as applied to claim 1 above and further in view of Kang (US 20230044057).
Regarding claim 2, Ellis in view of Logan discloses the method as claimed in claim 1, further comprising: acquiring a viewer's real-time chat for the first video transmitted in the real-time streaming manner (see include, but are not limited to, Ellis: figures 50-51, paragraphs 0106, 0216; DeWeese: figures 9, 15a-16, paragraphs 0058-0060); and
generating reaction text for at least a part of the real-time chat (generating reaction/chat text for at least a part of the real time chat using chat model- (see include, but are not limited to, Ellis: figures 50-51, paragraphs 0106, 0216; DeWeese: figures 9, 15a-16, paragraphs 0058-0060).
Ellis does not explicitly use the term “chatbot” model.
Kang discloses generating reaction text for at least a part of real-time chat using a chatbot model (see paragraph 0018, 0019, 0080-0083).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Ellis in view of Logan with the teaching of using chatbot model as taught by Kang in order to yield predictable result of improving conveniences for generating chat/comment in watching video content (see paragraphs 0004, 0018, 0029).
Regarding claim 3, Ellis in view of Logan and Kang discloses the method as claimed in claim 2, further comprising:
generating a second voice based on the text content and the reaction text (second voice associated based on the text content for different portion of content/reaction – see include, but are not limited to, Logan: paragraphs 0031, 0054, 0187) ;
generating a second view based on at least one of the text content or the second voice (see include, but are not limited to, Logan: paragraphs 0031, 0054, 0187; Kang: figures 6-8); and
transmitting a second video comprising the second voice and the second view in the real- time streaming manner (see include, but are not limited to, Logan: paragraphs 0031, 0054, 0187; Kang: figures 6-8).
Regarding claim 4, Ellis in view of Logan and Kang discloses the method as claimed in claim 2, wherein the generating of the reaction text for at least a part of the real-time chat comprises:
generating, using the chatbot model, reaction text associated with broadcast content for at least a part of the real-time chat, based on the broadcast outline or the text content (see include, but are not limited to, DeWeese: figures 11, 15A-16; Kang: paragraphs 0080-0083, 0111-0113).
Regarding claim 5, Ellis in view of Logan and Kang discloses the method as claimed in claim 2, wherein the generating of the reaction text for at least a part of the real-time chat comprises: selecting one or more real-time chats associated with broadcast content from the real- time chat using a similarity measurement model (e.g., comments/chats from users using a particular chat group - see include, but are not limited to, DeWeese: figures 11, 15A-16; Kang: paragraphs 0078, 0085, 0111-0113 ); and
generating the reaction text for the selected one or more real-time chats using the chatbot model (see include, but are not limited to, DeWeese: figures 11, 15A-16; Kang: paragraphs 0078, 0085, 0111-0113).
Regarding claim 6, Ellis in view of Logan discloses the method as claimed in claim 2, further comprising: filtering the real-time chat using a detection model configured to determine of input text (filtering real-time chat using detecting model/software configured to determine input text/comments for particular group chat, preferred or popular comments, topic, positive comment, negative comment – see include, but are not limited to, DeWeese: figures 1D, 11, 15A-16, 20, paragraph 0135; Logan: paragraphs 0054, 0222, Kang: paragraphs 0023, 0024, 0081, 0084-0086, 0111). However, Ellis in view of Logan and Kang does not explicitly disclose filtering chat using a hate-speech detection model to determine harmfulness of input text. Official Notice is taken that filtering real time chat using a hat-speech model detection model configured to determine harmfulness of input text to filtering inappropriate language, violent, hateful or undesired comment/chat using a detection model/software configured to determine violence, inappropriate or hateful language of input text of chat/comment/post is well-known in the art. For example, filtering comments/chat with hateful, violent or rating/curses that are not inappropriate for minor, children or user based on user profile/preference.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the well-known teaching in the art of filtering chat/comments using detection model configured to determine harmfulness of input text in order to yield predictable result of filtering undesired/inappropriate text from displaying to user such as children or anyone who does not want to receive hate-speech language.
See also Harb et al. (US 20190191209) for teaching of filtering user-generating content before displaying to other users based on user preferences (see paragraph 0002).
Regarding claim 7, Ellis in view of Logan and Kang discloses the method as claimed in claim 2, wherein the viewer's real-time chat comprises at least one of a text chat, an image chat, a sound chat, or a video chat (see include, but are not limited to, Ellis: figures 50-51; DeWeese: figures 12, 15-16, paragraph 0058, 0091).
Claims 8-9 are rejected under 35 U.S.C. 103 as being unpatentable over Ellis et al. (US 20150312640) in view of Longan et al. (US 20190206128) and Kang (US 20230044057) as applied to claim 3 above and further in view of Pham et al. (US 10897637).
Regarding claim 8, Ellis in view of Logan discloses the method as claimed in claim 3, wherein the generating of the second voice based on the text content and the reaction text comprises: generating the second voice using a transmission priority associated with the text content (using transmission priority associated with keyword, closed captioning of the selected content) and a transmission information associated with the reaction text (see for example, Ellis: figures 1a, 30-31; Logan: paragraphs 0031, 0054, 0185, 0222). However, Ellis in view of Logan and Kang does not explicitly disclose the information comprises transmission priority assorted with the reaction text.
Pham discloses generating audio/output using a transmission priority associated with text content and a transmission priority associated with reaction text (e.g., selecting of chat of most relevant content stream from multiple stream sources such as donation text from stream A or stream C and reduce volume of chats from less relevant content stream (see include, but are not limited to, figure 6, col. 11, lines 30-49).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Ellis in view of Logan and Kang with the teaching of transmission priority of reaction text as taught by Pham in order to yield predictable result of filtering text based on desired type of text based on user selection.
Regarding claim 9, Ellis in view of Logan, Kang and Pham discloses the method as claimed in claim 8, wherein the viewer's real-time chat comprises a donation chat associated with a donation (donation text 612 from stream C) and a general chat (text input of stream B) not associated with a donation, the transmission priority associated with the reaction text precedes the transmission priority associated with the text content (providing donation text of stream C and not providing input text of stream B), and among the reaction text, a transmission priority associated with reaction text for the donation chat precedes a transmission priority associated with reaction text for the general chat (see include, but are not limited to, Pham: figure 6, col. 11, lines 30-49).
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Ellis et al. (US 20150312640) in view of Longan et al. (US 20190206128) as applied to claim 1 and further in view of Yang et al. (US 20220239988).
Regarding claim 14, Ellis in view of Logan discloses the method as claimed in claim 1, wherein the generating of the text content based on the broadcast outline comprises generating the text content based on a broadcast outline in an image format (generating of text content such as description, closed captioning text using model based on broadcast outline/information in broadcast content – see include, but are not limited to, Ellis: paragraph 0129, 0178, 0183). However, Ellis does not explicitly disclose generating using an image recognition model the text based on broadcast outline in an image format.
Yang discloses generating, using an image recognition model, the text content based on a broadcast outline in an image format (generating, using an image recognition model, a text keyword, based on a broadcast outline in an image format with scene keyword, first item keyword, etc. – see include, but are not limited to, figures 6, 9, 11, paragraphs 0089-0090, 0094).
Therefore, it would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed invention to modify Ellis in view of Logan with the teaching of generating using an image recognition, the text content as taught by Yang in order to yield predictable result of improving efficiency of information interaction and user experience in recognizing text in image (see paragraphs 0006, 0091).
Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over Ellis et al. (US 20150312640) in view of Longan et al. (US 20190206128) as applied to claim 1 and further in view of Yang et al. (US 20220239988) or Harb et al. (US 20190191209).
Regarding claim 16, Ellis in view of Logan discloses the method as claimed in claim 1, wherein the generating of the text content based on the broadcast outline comprises generating the text content based on a broadcast outline in a video format (e.g., generating text such as subtitle, closed captioning, etc. for displaying to retrieve text or closed captioning, subtitle of content from video content – see include, but are not limited to, - see include, but are not limited to, – see include, but are not limited to, Ellis: paragraphs 0129, 0143, Longan: paragraphs 0031, 0053-0054, 0222).
However, Ellis in view of Logan does not explicitly disclose generating using a video recognition model.
Yang or Harb discloses generating, using a video recognition model, the text content based on a broadcast outline in a video format (generating, using an image(still video) or video recognition model, a text/keyword, based on a broadcast outline in an image/video format – see include, but are not limited to, Yang: figures 6, 9, 11, paragraphs 0089-0090, 0094; Harb: paragraphs 0011, 0062).
Therefore, it would have been obvious to one of ordinary skilled in the art before the effective filing date of the claimed invention to modify Ellis in view of Logan with the teaching of generating using a video recognition, the text content in video format as taught by Yang/Harb in order to yield predictable result of improving efficiency of information interaction and user experience in recognizing text in image (see Yang: paragraphs 0006, 0091; Harb: paragraph 0062).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Ziro Jr. (US 20160021412) discloses multi-media presentation system for presenting different video with difference voices to users based on user reaction (see for example, paragraphs 0069, 0078-0079, 0086, 0101-0103).
Thomas et al. (US 20180288490) discloses systems and methods for navigating media assets and reaction text from users while viewing video content.
Kumar et al. (US 20210029406) discloses systems and methods for applying behavioral based parental controls for media assets.
Yoshida et al. (US 20220103873) discloses providing avatar for distributor and control movement and audio output for avatar based on movement and speech of the distributor.
Quisenberry et al. (US 20220109911) discloses methods and apparatus for determining aggregate sentiments and displaying aggregated sentiments.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AN SON PHI HUYNH whose telephone number is (571)272-7295. The examiner can normally be reached 9:00 am-6:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, ANHTUAN T. NGUYEN can be reached at 571-272-4963. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AN SON P HUYNH/Primary Examiner, Art Unit 3795
July 20, 2026