Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 2, 3, 7-20 rejected under 35 U.S.C. 103 as being unpatentable over Kakoyainnis: 20200013380 hereinafter Ka further in view of Merler: Automatic Curation of Sports Highlights Using Multimodal Excitement Features (copy provided by Examiner; copyright 2018 and hereinafter Merl).
Regarding claim 2
Ka teaches:
A method (Ka: ¶ 67; Fig 5: such as practiced upon a computer, computer system by instantiation and/or execution of coded instructions from memory), comprising: receiving one or more media content items (Ka: Fig 5-7: computer system 102 receives an audio track, processes same, provides an audio track to a user device such as that of figure 6 which additionally receives, processes, and provides the audio track);
using machine learning to identify one or more audio segments of interest in each of the one or more media content items, based at least in part on an analysis of content included in a corresponding media content item, wherein each of the identified audio segments is associated with one or more automatically determined tags (Ka: Abstract; ¶ 21, 59, 64, 72-76, 87, 130-133, etc.; Figs 5-7: system operates to segment an audio track corresponding to a particular one of the one or more media content items such as by employ of machine learning computer algorithms to enrich the corresponding media content item such as with a textual element such as a keyword, tag, descriptive copy and/or title automatically derived from the item, segment(s) thereof);
generating a video clip (Ka: Abstract, ¶ 59-61, 83, 84, 139, 144-147 etc.: system for automated production of video loops from images and videos identifies or otherwise recommends audio segments for pairing with video, sequences of images etc. for delivery to a user such as for playback upon a media player) for a specific user (Ka: ¶ 62, 83, 84, 127, 151: artifact created for and delivered to a specific user such as based on user preferences, interests, etc. based on matching user information, tags, keywords, etc.); that comprises a plurality of video segments (Ka: ¶ 62, 83, 84, 88, 105, 127: system compiles two or more constituent segments into one clip, loop, etc.) extracted from media content (Ka: ¶ 80-84, 127, 130: system determines topical audio segments pairs same with visual asset media in data storage) said content comprising such as a from a plurality of episodes of media content items (Ka: ¶ 88, 106, 119: segments for combining into a loop comprise selections from one or more shows, episodes thereof with respect to topics, themes, etc.), including:
using machine learning to automatically select, recommended audio segments from the identified audio segments based at least in part on one or more prior user interactions of a user and the automatically determined tags of the identified audio segments (Ka: Abstract, ¶ 59-61, 64, 83, 139 etc.: such as by utilizing an algorithm for constructing a video clip based on a user interest or listening history and additionally based on pairing of segments of audio and video such as based on tags, keywords, topics, etc.), and
identifying, for inclusion in the video clip, video segments of the one or more media content items that correspond to the recommended audio segments (Ka: Abstract, ¶ 59-61, 83, 139 etc.; Figs 5-7: system identifies segments for pairing with video, sequences of images, etc. for delivery to a user such as for playback upon a media player); and
providing the video clip for playback on a device associated with a user (id.).
Ka does not explicitly teach using a machine learning recommender to winnow particular video segments of media content items used to generate a singular video clip that comprises a plurality of video segments extracted from different episodes of the media content items, thereby selecting audio segments from identified audio segments to pair with video segments extracted from different episodes of the media content items.
In a related field of endeavor Merl teaches a system and method for automatically extracting highlight clips (Merl: Abstract) comprising receiving one or more media content items (§ III; pp 4, 6: system receives a plurality of video feeds as input);
using machine learning to identify one or more audio segments of interest in each of the one or more media content items based at least in part on an analysis of content included in a corresponding media content item (Merl: § III; pp 4, 6: system determines audio segments of interest using classifiers to determine crow cheering and other indicia of excitement in audio corresponding to one or more video segments which are determined to be highlight segments based thereon),
wherein each of the identified audio segments is associated with one or more automatically determined tags (Merl: § III pp 6; Fig 2: system determines metadata of highlight segments such as player names, locations, scoring information therein);
generating a video clip for a specific user (Merl: § 1, pp2; § 2, pp3: highlight clips generated based on preferences of a viewer such as , player, location, etc.)
that comprises a plurality of video segments extracted from different episodes of the media content items (Merl: Abstract; system produces highlights summarizing more exciting moments in a content series such as games, rounds, etc. of a match, tournament, etc.; said highlights extracted from distinct segments of distinct temporally and location separated units of the content series); including:
using machine learning to automatically select, for the specific user, recommended audio segments from the identified audio segments based at least in part on one or more prior user interactions of the specific user and the automatically determined tags of the identified audio segments (Merl: § I, pp 2; § III, p 6: system performs personalized highlight extraction utilizing various machine learning algorithms to personalize selection of highlight clips such as based on a user’s interaction revealing explicit preferences operable to resolve metadata parameters; such as in response to a user assertion to “show me all highlights of player X at hole Y during the tournament”); and
identifying, for inclusion in the video clip, video segments of the one or more media content items that correspond to the recommended audio segments (Merl: Abstract; § III, pp 6, 7: system determines audio of interest, metadata thereof and matches same to a video clip determined to comprise a highlight); and
providing the video clip for playback on a device associated with the specific user (Merl: § pp 2; Fig 1, 2: clips shared in a user interface for user designated playback of video highlights therein. It would have been obvious to one of ordinary skill in the art before the effective filing date of the instant application to utilize the Merl taught curation and provision of highlight videos in a clip based manner to provide topically generated videos to a specific user(s) based on the Ka system and method for at least the purpose of targeting or recommending a video more likely to be ingested or otherwise desired by a user; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 3
Ka in view of Merl teaches or suggests:
The method of claim 2, wherein the one or more prior user interactions of the specific user includes one or more other interactions by the specific user with respective audio segments in the media content items (Ka: ¶ 64, 106, 108: etc.: system improves based on prior user actions with respect to the system which learns from the curator); (Merl: § II, pp 3: user interacts to personalize the determination of particular highlights, output thereof). The claim is considered obvious over Ka as modified by Merl as addressed in the base claim as it would have been obvious to apply the further teaching of Ka and/or Merl to the modified device of Ka and Merl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 7
Ka in view of Merl teaches or suggests:
The method of claim 3, further comprising: updating a machine learning model based on an interaction by the specific user, the machine learning model configured to identify respective audio segments to recommend (Ka: ¶ 8, 42, 55, 64, 72, 106, 108: system trained based on user behavior with respect to audio segments, training curated by a user; updated in such a way that the system improves based on prior user actions with respect to the system which learns from the curator); (Merl: § III, pp 4, 5: system iteratively refined updated based on previous data; including previous interaction data such that weights “for each component are learned via cross-validation,”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the instant application to regularly update the machine learning model based on the availability of new content and to include the taught user interaction features to winnow such new content. The claim is thus considered obvious over Ka as modified by Merl as addressed in the base claim as it would have been obvious to apply the further teaching of Ka and/or Merl to the modified device of Ka and Merl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 8
Ka in view of Merl teaches or suggests:
The method of claim 2, wherein: the one or more prior user interactions of the specific user includes one or more interactions indicating types of content that the specific user is interested in, the one or more interactions selected from the group consisting of: the specific user listening to a particular audio segment to completion; the specific user sharing the particular audio segment with other users of different electronic devices; and the specific user subscribing to the content corresponding to the particular audio segment (Ka: ¶ 59, 68, 72, 139: such as by tracking a user listening history, user media sharing, subscription to a media item such as a podcast); (Merl: § I, pp 3: system tracks user viewing to completion). The claim is considered obvious over Ka as modified by Merl as addressed in the base claim as it would have been obvious to apply the further teaching of Ka and/or Merl to the modified device of Ka and Merl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 9
Ka in view of Merl teaches or suggests:
The method of claim 2, further comprising: receiving an indication of a user action associated with advancing an active audio segment; selecting for the specific user a second recommended audio segment from at least the identified audio segments; and automatically providing the second recommended audio segment (Ka: Abstract; ¶ 88, 149; Fig 13A: such as by tracking user operation to manipulate a media and generating autoplay media from disparate tracks based thereon); (Merl: § III, pp 5, 6: highlight reel presents successive clips, such as in a ranked order upon an interactive dashboard operate that advancing through clips proceeds through clips in the ranked order). The claim is considered obvious over Ka as modified by Merl as addressed in the base claim as it would have been obvious to apply the further teaching of Ka and/or Merl to the modified device of Ka and Merl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 10
Ka in view of Merl teaches or suggests:
The method of claim 2, further comprising: receiving an indication of a user action associated with an active audio segment; and providing a full corresponding media content item associated with the active audio segment (Ka: ¶ 88, 106: a user experiencing a clip may expand out the clip to a full episode related thereto); (Merl: § III, pp 3: highlights reference back to source , match, episode, etc. thereby enabling retrieval thereof). The claim is considered obvious over Ka as modified by Merl as addressed in the base claim as it would have been obvious to apply the further teaching of Ka and/or Merl to the modified device of Ka and Merl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 11
Ka in view of Merl teaches or suggests:
The method of claim 2, wherein the analysis of content included in the corresponding media content item includes identifying topics associated with identified word content in the corresponding media content item (Ka: ¶ 73, 74, 88, etc.: system operates to reify topics by keyword, term, etc.); (Merl: ¶ 111, pp 5: identified word, expression, etc. content used to determine excitement level, clip timings based thereon). The claim is considered obvious over Ka as modified by Merl as addressed in the base claim as it would have been obvious to apply the further teaching of Ka and/or Merl to the modified device of Ka and Merl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 12
Ka in view of Merl teaches or suggests:
The method of claim 11, wherein the automatically determined tags of the identified audio segments are based on the identified topics (Ka: ¶ 73-76, 88, 105, 106: “audio segments are tagged with relevant descriptors,” in this way tags are conflated with topic and a keyword may be considered a tag which reifies a topic); (Merl: § III, pp 4; Fig 2: system derives metadata from keywords identified within video segments, attaches said metadata to the relevant segment). The claim is considered obvious over Ka as modified by Merl as addressed in the base claim as it would have been obvious to apply the further teaching of Ka and/or Merl to the modified device of Ka and Merl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 13
Ka in view of Merl teaches or suggests:
The method of claim 2, wherein the analysis of content included in the corresponding media content item includes automatically transcribing each of the media content items (Ka: ¶ 7, 19, 73-76, etc.: such as by conversion of talk-based audio to text, tag associated therewith such as based on raw speech data, text data, or voice recognition data from the audio track); (Merl: § I, pp1, § III, pp 4: system models a commentators tone, extracts expressions using a speech to text service). The claim is considered obvious over Ka as modified by Merl as addressed in the base claim as it would have been obvious to apply the further teaching of Ka and/or Merl to the modified device of Ka and Merl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 14
Ka in view of Merl teaches or suggests:
The method of claim 13, wherein transcribing each of the media content items includes automatically identifying one or more speakers of content in each of the media content items (Ka: ¶ 73, 79, 81, etc.: speech processed with a neural network, voice recognition module, etc.); (Merl: § I, pp1, § III, pp 4: system identifies a commentator, models commentator tone, extracts expressions using a speech to text service). Examiner has taken official notice which Applicant has failed to timely and explicitly traverse; it is thus accepted as Admitted Prior Art (APA; please see MPEP 2144.03) that speaker recognition would have comprised an obvious inclusion for at least the purpose of diarizing a text based media comprising plural speakers. The claim is thus considered obvious over Ka as modified by Merl as addressed in the base claim as it would have been obvious to apply the further teaching of Ka and/or Merl to the modified device of Ka and Merl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 15
Ka in view of Merl teaches or suggests:
The method of claim 2, wherein the analysis of content included in the corresponding media content item includes automatically identifying advertisements in the corresponding media content item (Ka: ¶ 1, 2, 4, 59, 121-124: such as by locating advertising based on determined topics, metadata, etc. of particular media); (Merl: . The claim is considered obvious over Ka as modified by Merl as addressed in the base claim as it would have been obvious to apply the further teaching of Ka and/or Merl to the modified device of Ka and Merl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 16
Ka in view of Merl teaches or suggests:
The method of claim 2, wherein the analysis of content included in the corresponding media content item includes automatically identifying music in the corresponding media content item (Merl: § III, pp 4, § V, pp 8: system recognizes music, delineates same from speech. The claim is considered obvious over Ka as modified by Merl as addressed in the base claim as it would have been obvious to apply the further teaching of Ka and/or Merl to the modified device of Ka and Merl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 17
Ka in view of Merl teaches or suggests:
The method of claim 2, wherein the analysis of content included in the corresponding media content item includes automatically identifying questions in the corresponding media content item (Ka: ¶ 73-75: system segments text over phonemes); (Merl: § III, pp 4: system trains sub-classifiers to identify specific words, phrases). Examiner has taken official notice which Applicant has failed to timely and explicitly traverse; it is thus accepted as Admitted Prior Art (APA; please see MPEP 2144.03) that parsing a text for questions and answers by automatic identification of same would have comprised an obvious inclusion for at least the purpose of identifying a discourse structure, determine questions relevant to resolve images or videos, classify speech acts with respect to additional labels for determining images or videos, identify questions relevant to particular topics, segments ascribed thereto, etc. The claim is thus considered obvious over Ka as modified by Merl as addressed in the base claim as it would have been obvious to apply the further teaching of Ka and/or Merl to the modified device of Ka and Merl; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 18
Ka in view of Merl teaches or suggests:
The method of claim 2, wherein the video clip includes one or more visual indicators of a particular audio segment (Ka: Abstract; ¶ 11, 149, etc.; Fig 2, 13C: system pairs visual assets to identified content, such as by inclusion of a textual element, asset tag or other meta-, auxiliary, etc. data in the resulting container such as for display to a user or consumer of the media); (Merl:§ II, pp 3; § III, pp 6: system overlays name, metadata, keywords etc. within a media item), the one or more visual indicators selected from the group consisting of:
a name of the corresponding media content item (Ka: ¶ 21, 80, etc.: textual elements comprise title associated with an audio elements, segment, etc.); (Merl:§ II, pp 3; § III, pp 6: system overlays name, metadata, keywords etc. within a media item); speakers in the corresponding media content item (please see claim 14 supra diarizing speakers in a media considered well-known); subtitles and speaker information in the particular audio segment (Merl:§ II, pp 3; § III, pp 6: system overlays name, metadata, keywords etc. within a media item); tags corresponding to the particular audio segment (Ka: ¶ 21, 80, etc.: textual elements comprise metatags, keywords associated with the audio elements, segment, etc.); references to the corresponding media content item; and references to content applications for playing the corresponding media content item. While Ka in view of Merl do not explicitly discuss the inclusion of every item in the group the remaining items are considered obvious as a matter of design choice on the part of an implementor of the system and/or of a creator of media therein; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 19, 20—the claims are considered to recite substantially similar subject matter to that of claim 2 and are similarly rejected.
Claims 4-6 rejected under 35 U.S.C. 103 as being unpatentable over Kakoyainnis: 20200013380 hereinafter Ka further in view of Merler: Automatic Curation of Sports Highlights Using Multimodal Excitement Features (copy provided by Examiner; copyright 2018 and hereinafter Merl) as applied to claims 2, 3, 7-20 supra, and further in view of Chawla: 11451598 hereinafter Chaw .
Regarding claim 4
Ka in view of Merl teaches or suggests:
The method of claim 2, wherein: selecting, for the specific user, the recommended audio segments from the identified audio segments based at least in part on the one or more prior user interactions of the specific user (Ka: ¶ 64, 106, 108: etc.: system improves based on prior user actions with respect to the system which learns from the curator); (Merl: § II, pp 3: user interacts to personalize the determination of particular highlights, output thereof) but does not teach identifying types of media content that the specific user is not interested in based on interactions in which the specific user skips or ignores particular media content.
In a related field of endeavor Chaw teaches a system for generating a preview segment of a recommended media comprising receiving one or more media content items (Chaw: Abstract; Col 1:46-1:67: system identifies one or more media items, segments thereof, to recommend to a user and gauges user interest in same);
using machine learning to identify one or more audio segments of interest in each of the one or more media content items based at least in part on an analysis of content included in a corresponding media content item, wherein each of the identified audio segments is associated with one or more automatically determined tags (Chaw: 2:20-2:27, 7:17-7:27, 8:53-8:64, etc.: system identifies segments, media items, etc. of potential interest for presentation, recommendation, etc. to a user; said segments, etc. based on tag data such as genre, musical qualities and selected using a machine learning system);
generating a video clip for a specific user (Chaw: 1:46-1:67; 2:20-2:27, 8:43-8:49, 8:60-9:7: system determines segments to present to a particular user based on user specific factors including, user media consumption, user metadata preferences, etc. similarity of same to other user consumption, preferences, etc.) including identifying types of media content that the specific user is not interested in based on interactions in which the specific user skips or ignores particular media content (Chaw: Col 1:57-1:67, 2:16-2:27, 8:10-8:25: user interactions in the form of “watch” data incorporated into the machine learning recommender; said data including a user choosing to skip the preview or ignore the preview in favor of deferment).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the instant application to utilize the Chaw taught user specific user play dynamics to provide the Ka in view of Merl topically generated videos to a specific user(s) for at least the purpose of targeting or recommending a video more likely to be ingested by a user; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 5
Ka in view of Merl in view of Chaw teaches or suggests:
The method of claim 4, wherein the one or more prior user interactions of the specific user includes a swiping gesture causing a next or previous audio segment to be played back before completion of playback of the recommended audio segment (Ka: ¶ 64, 106, 108: etc.: system improves based on prior user actions with respect to the system which learns from the curator); (Merl: § II, pp 3: user interacts to personalize the determination of particular highlights, output thereof); (Chaw: Col 1:57-1:67, 2:16-2:27, 8:10-8:25; Fig 6: user interactions in the form of “watch” data incorporated into the machine learning recommender; said data including a user choosing to skip the preview or ignore the preview in favor of deferment; see particularly figure 6 which embodies a swipe type media interface). Examiner has taken official notice which Applicant has failed to timely and explicitly traverse; it is thus accepted as Admitted Prior Art (APA; please see MPEP 2144.03) that the recited user interface features would have comprised an obvious inclusion for at least the purpose of providing a user aspects of trick play to thereby deliver an expected media playback experience. The claim is thus considered obvious over Ka as modified by Merl, and Chaw as addressed in the base claim as it would have been obvious to apply the further teaching of Ka, Merl, and/or Chaw to the modified device of Ka, Merl, and Chaw; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 6
Ka in view of Merl in view of Chaw teaches or suggests:
The method of claim 5, wherein: the one or more prior user interactions of the specific user includes a second swiping gesture causing additional content about the next or previous audio segment to be displayed (Chaw: Col 1:57-1:67, 2:16-2:27, 8:10-8:25; Fig 6: user interactions in the form of “watch” data incorporated into the machine learning recommender; said data including a user choosing to skip the preview or ignore the preview in favor of deferment; see particularly figure 6 which embodies a swipe type media interface). Examiner has taken official notice which Applicant has failed to timely and explicitly traverse; it is thus accepted as Admitted Prior Art (APA; please see MPEP 2144.03) that the recited user interface features would have comprised an obvious inclusion for at least the purpose of providing a user aspects of trick play to thereby deliver an expected media playback experience; one of ordinary skill in the art would have expected only predictable results therefrom. The claim is thus considered obvious over Ka as modified by Merl, and Chaw as addressed in the base claim as it would have been obvious to apply the further teaching of Ka, Merl, and/or Chaw to the modified device of Ka, Merl, and Chaw; one of ordinary skill in the art would have expected only predictable results therefrom.
Response to Arguments
Applicant’s arguments in concert with amended claims, see Remarks and Claims, filed 5/26/26, with respect to the rejection(s) of claim(s) 2-20 under 35 USC 103 over Kakoyainnis in view of Chawla have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Kakoyainnis in view of Merler; and Kakoyainnis in view of Merler in view of Chawla.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PAUL C MCCORD whose telephone number is (571)270-3701. The examiner can normally be reached 730-630 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, CAROLYN EDWARDS can be reached at (571) 270-7136. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PAUL C MCCORD/ Primary Examiner, Art Unit 2692