DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
All objections/rejections not mentioned in this Office Action have been withdrawn by the Examiner.
Status of the Claims
Prior to entry of the amendment(s) and/or consideration of the argument(s), the status of the claims is as follows.
Claim(s) 1-20 is/are pending.
Claim(s) 1-6, 12-15 and 20 is/are rejected under 35 U.S.C. 102(a)(1) and 102(a)(2) as being anticipated by Yun (U.S. Pat. No. 9,430,115, hereinafter Yun).
Claim 7-10, and 16-18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yun as applied to claims 1 and 12 above, and further in view of Mahyar (U.S. Pat. No. 10999566, hereinafter Mahyar).
Claims 11 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yun and Mahyar as applied to claim 7 and 16 above, and further in view of Non-patent Literature to Eisenberg (Eisenberg, J & Finlayson, M. (2021) 'Narrative Boundaries Annotation Guide', Journal of Cultural Analytics. 6(4) https://doi.org/10.22148/001c.30698, hereinafter Eisenberg).
Response to Amendments
Applicant’s amendment filed on 26 May 2026 has been entered.
In view of the amendment to the claim(s), the amendment of claim(s) 1-7 and 9-20; the cancellation of claim(s) 8; and the addition of claim(s) 21 have been acknowledged and entered.
In view of the amendment to claim(s) 2 and 13, the objection to claim(s) 2 and 13 is withdrawn.
In view of the amendment to claim(s) 1-7 and 9-20 and the cancellation of claim(s) 8, the rejection of claims 1-20 under 35 U.S.C. §102 and 103 as previously presented is withdrawn.
In light of the amended/newly added claims, new grounds for rejection under 35 U.S.C. §102, 35 U.S.C. §103 and 35 U.S.C. §112 are provided in the action below.
Response to Arguments
Applicant’s arguments regarding the prior art rejections under 35 U.S.C. §102/103, see pages 13-17 of the Response to Non-Final Office Action dated 25 February 2026, which was received on 26 May 2026 (hereinafter Response and Office Action, respectively), have been fully considered.
With respect to the rejection(s) of claim(s) 1, 12, and 20 under 35 U.S.C. §102(a)(1) and 102(a)(2) as being anticipated by Yun, applicant asserts that Yun fails to teach or suggest all limitations of claims 1, 12, and 20 as amended. Applicant’s arguments, at least with respect to the use of embeddings is persuasive. As such, the rejection of claims 1, 12, and 20 under 35 U.S.C. §102 are withdrawn.
Claim 8 is canceled in the response. Therefore, the rejection of claim 8 is withdrawn.
Applicant further argues that the rejection(s) of dependent claims 2-7, 9-11, and 13-19 should be withdrawn for at least the same reasons as independent claims 1, 12, and 20. Applicant’s arguments in light of the amended claims are persuasive. As such, the rejections of claims 2-7, 9-11, and 13-19 under 35 U.S.C. §102 and 35 U.S.C. §103 are withdrawn.
However, upon further consideration, new ground(s) of rejection under 35 U.S.C. §103 are made in light of combinations of Yun, Mahyar, and Eisenberg.
The Applicant has not provided any further statement and therefore, the Examiner directs the Applicant to the below rationale.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 1-7 and 9-21 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
Regarding claims 1, and mutatis mutandis claims 12 and 20, the “determining a different position for the text data within the audio-derived text sequence, wherein the different position is associated with the second time in the timeline” is not clearly taught in the specification. The above limitation is recited in claim 1 requires a determination of a position in a timeline, where the position is associated with the second time. Respectfully, the scope of at least the “second time,” lacks reasonable boundaries such that it encompasses a boundless genus of temporal relationships that are not supported by the original disclosure.
In performing this analysis, it is noted that the vast majority of the amendments provided are not expressly recited in the specification. As such, support for said amendments must be determined by analysis and interpretation of the specification and claims as filed. The specification as filed does not disclose a “second time,” or a definition of said second time. As such, we look to usage in claim 1 to understand the broadest reasonable interpretation of “second time”. Claim 1, as amended, tethers the “second time” to the claim only by correspondence of “a content of the segment… in the timeline” to the second time. The second time itself remains undefined in the claim, as the second time is defined only as a passive recipient of a relationship. In both logic and claim construction, the mere fact that Element A(Content) “corresponds to” Element B(second time), does not mandate a bijective (one-to-one and reciprocal) relationship such that Element B corresponds to Element A. As such, the second time, as recited in amended claim 1, is lacking independent structural and chronological boundaries. Further, it is noted that “the second time in the timeline corresponding to the content of the segment” is later included as a condition (i.e., “based on…”) for the determining, said condition appears to indicate that correspondence between the two is considered. However, as the correlation is not recited in the claims, the basis on said correlation is unclear. Claim 1 further defines the “segment” as a portion of “the content item” which is “corresponding to the deviation in the timeline.” The “deviation in the timeline” is not specifically defined in the specification, but the specification does provide examples such that one skilled in the art would understand a deviation as referring to a flashback, flashforward, or a recap. (See Instant Application, [0083]). It is noted that “corresponding” establishes a relationship (e.g., similar, connected, matching, or equivalent to) but does not require a direct connection. Further, and as above, the “segment of the content item” is corresponding to “the deviation in the timeline”, thus the segment has a relation to the “deviation in the timeline”, but the segment is not necessarily the “deviation in the timeline”. As such, and in light of the above analysis, a second time broadly refers to any time in a timeline to which a portion of a content item has a relationship, where the portion has a relationship to a deviation in the timeline.
Regarding the interaction between the second time and the different position, the specification fails to establish such a relationship. The specification fails to recite the word “position” or the phrase “different position”. Though the specification does reference the concept of positions indirectly, such as through words like group or organize, it is unclear how such phrases could be interpreted to support determining a specific “different position” in a timeline based on a second time as described above. The “different position” also lacks specification support and is only limited insofar as the second time is limited. The further limitation of “based on the second time in the timeline corresponding to the content of the segment” asserts a basis which is never established to actually exist in the claim. Stated more specifically, the claim does not establish that the second time corresponds to the content of the segment, thus the meaning of being based on that asserted fact fails to provide a clear limit on either the different position or the mechanism of determination said different position. The further limitation of “based on…a position of the text data within a temporal sequence” relies on necessary components (an object having a position such that a new position may be described as different) and disconnected claim parts which lack clear specification support and/or further description. The limitation recites “a temporal sequence of the audio-derived text sequence” which, though similar to the previously recited “temporal location”, is not connected to said “temporal location” or to said “first time”, and does not have clear specification support. As such, the broadest reasonable interpretation of the “different position” is any position in the audio-derived text sequence which can be associated with the second time.
Respectfully, the cited paragraphs and the specification generally, fails to teach or suggest “determining a different position for the text data within the audio-derived text sequence” based on a second time as explained above. The specification does not demonstrate possession of the broad set of temporal relationships encompassed by the “second time” or any relationship between a time which may be a second time the subsequent determination of any particular position in the timeline. The specification further fails to clearly recite or explain any mechanism for how a specific single time may be used or applied to determine any particular position for selected text data in a timeline. Therefore, claims 1, 12, and 20 contain limitations which lack specification support and are rejected.
Regarding claims 2 and 14, the limitation “wherein the segment comprises… content replay… [and] content preview” lack specification support. The segment refers to a portion of a content item corresponding to a deviation in the timeline. The cited portions of the specification do not provide clear support for a segment comprising a content replay or a content preview. Regarding a content preview, the phrase “content preview” does not occur in the specification or in the claims as filed. The specification does disclose the word “preview,” with regards to generating previews. (see Instant Application, ¶ [0024], [0029], [0049], [0060], [0063], [0080], [0120]). However, said previews are not described as part of the content item, but as a result derived from processing the content item. The content item and segments thereof are not described as including a preview. representative example of The phrase “content replay” and/or the word “replay” does not occur in the specification and clear support for a “content replay”, outside of a “recap”, could not be found. Therefore, claims 2 and 14 contain limitations which lack specification support and are rejected.
Regarding claims 2-7, 9-11, and 13-21, claims 2-7, 9-11, and 13-21 depend from claims 1, 12, and 20 and incorporate all limitations therefrom. Therefore, claims 2-7, 9-11, and 13-21 are rejected for at least the same reasons as claims 1, 12, and 20.
Regarding claims 1-7 and 9-21, examiner has attempted to indicate all known portion of the newly amended claims which lack specification support. However, it is noted that numerous parts of the amended claims are internally reliant and future amendments may result in later rejections under 35 USC 112(a) and/or 112(b). Applicant is advised to review the remaining limitations of the claims as amended, both now and in response to this rejection, to confirm specification support of the limitations and avoid unnecessary further rejection.
Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 1-7, 9-10, 12-18 and 20-21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yun in view of Mahyar.
Regarding claim 1, Yun discloses A system comprising: memory; and one or more processors coupled to the memory and configured to perform operations comprising (Systems and methods described with relation to “ generating storylines associated with content” which is implemented using a “device 104” through “one or more processors 106” the processors being “configured to access and execute at least in part instructions stored in the one or more memories 108.”; Yun, ¶ Col. 2, line 64 - col 3, line 33): detecting a deviation in a timeline of a content item and a segment of the content item corresponding to the deviation in the timeline, (Discloses “generat[ing] one or more tags 132 associated with the content 124 based at least in part on the determined object data and the one or more occurrence times 212,” where the one or more tags can each indicate a time frame in the content (e.g., a tag for “a saddle appears at times 51:13 to 55:21.”) where the tags are presented as part of a storyline, and the storyline is defined by “associating tags 132 with similar or the same information across time in one or more pieces of the content 124.”; Yun, ¶ Col. 4, lines 41-46; Col. 14, lines 11-15) wherein a temporal location of the segment corresponds to a first time in the timeline and a content of the segment corresponds to a second time in the timeline (“the tags 132 provide machine-readable information about events 208 taking place at a particular point in the content 124” and “may include reference information, associating the tag with a particular portion of the content 124 or designating the particular portion of the content at which the event described by the tag occurs,” where “each event 208 has at least one corresponding occurrence time which specifies the point or interval in the content at which the event takes place” and a “time” in the “chronology of the story” (hereinafter, a chronological time) where, in the context of a flashback, the tag indicates a temporal location of the segment of the content corresponding to a first time, based on surrounding, non-contiguous events, and the event has a chronological time which is earlier in the story {a content of the segment corresponds to a second time in the timeline}; Yun, ¶ Col. 4, lines 34-40; col. 5, lines 15-23, and 59-67; Col. 15, lines 22-36); identifying text data, from an audio-derived text sequence associated with a content item, that corresponds to the segment of the content item (“determines text data associated with the one or more occurrence times 212 in the content 124,” where the text data is derived from the entire transcript or caption sequence {an audio derived text sequence}; Yun, ¶ Col. 14, lines 16-21), the audio-derived text sequence comprising at least one of a transcript of audio associated with the content item, captions of the audio associated with the content items and subtitles corresponding to the audio associated with the content item (“The text data may comprise information from retrieving one or more text captions, recognizing speech with a speech recognition module, or recognizing writing with an optical character recognition module” and may “be based at least in part on closed captions, open captions, subtitles, and so forth”; Yun, ¶ Col. 14, lines 16-25); based on the second time in the timeline corresponding to the content of the segment and a position of the text data within a temporal sequence of the audio-derived text sequence (Discloses “arranging” the tags “by increasing time of the chronology of the story” according to the chronological time of the events.; Yun, ¶ Col. 15, lines 22-36), determining a different position for the text data within the audio-derived text sequence, (The above arrangement “may result in a storyline which presents the scenes showing the flashbacks first, consistent with their earlier appearance in the internal chronology of the story as being earlier events,” which includes a different determined position for the text data within the storyline corresponding to the chronological occurrence of events within the content, where “The content 124 may comprise one or more of video data 202, audio data 204, [and] caption data 206 {audio derived text sequence}”; Yun, ¶ Col. 4, lines 60-63; Col. 15, lines 22-36) wherein the different position is associated with the second time in the timeline (The different position is “consistent with their earlier appearance in the internal chronology of the story as being earlier events,” where the chronological time is the second time in the “internal chronology {timeline}” where, in the context of “video data 202, audio data 204, [and] caption data 206” as the content, a different position is determined for each of these, and said position is associated with the second time.; Yun, ¶ Col. 4, lines 60-63; Col. 15, lines 22-36); determining a modified temporal sequence of the audio-derived text sequence (As above, the “storyline... presents the scenes showing the flashbacks first, consistent with their earlier appearance in the internal chronology of the story as being earlier events,” which includes a modified temporal sequence of the events (which necessarily includes a modified temporal sequence of the text sequence in the content, for each of “video data 202, audio data 204, [and] caption data 206”; Yun, ¶ Col. 4, lines 60-63; Col. 15, lines 22-36) based on the temporal sequence of the audio-derived text sequence and the different position within the audio-derived text sequence determined for the text data (Said modified temporal sequence is determined based on the content, where content can be any or all of “video data 202, audio data 204, [and] caption data 206”. As such, the temporal sequence can be based on the “internal chronology” of video, audio, caption (text), or any combination.; Yun, ¶ Col. 4, lines 60-63; Col. 15, lines 22-36); based on the audio-derived text sequence and the modified temporal sequence, generating a rearranged audio-derived text sequence associated with the content item (The “scenes” are presented {generated} “showing the flashbacks first, consistent with their earlier appearance in the internal chronology of the story as being earlier events,” as understood in the context of the caption data 206, this also includes the generation of the caption data rearranged “consistent with their earlier appearance in the internal chronology of the story”; Yun, ¶ Col. 15, lines 22-36), the rearranged audio-derived text sequence comprising the audio-derived text sequence rearranged according to the modified temporal sequence (“storyline 134 may be configured to present portions of the content designated by the plurality of tags 132 while omitting portions of the content 124 undesignated by the plurality of tags 132” which is understood to include not “omitting portions of the content 124 undesignated by the plurality of tags 132” and “a storyline 134 for a character which features may flashback and flash-forward scenes may result in a storyline which presents the scenes showing the flashbacks first, consistent with their earlier appearance in the internal chronology of the story” which is a rearrangement of the content {audio derived text sequence} according to the order of the event’s chronological appearance {the modified temporal sequence}; Yun, ¶ Col. 15, lines 22-36). However, Yun fails to expressly recite generating embeddings of the rearranged audio-derived text sequence.
Mahyar teaches “systems and methods to automatically generate textual descriptions for video content.” (Mahyar, ¶ Col. 2, lines 33-35). Regarding claim 1, Mahyar teaches generating embeddings of the rearranged audio-derived text sequence (“The second neural network may analyze the individual textual descriptions, aggregate and combine the individual textual descriptions, incorporate analysis of the audio, and output a textual description for the scene and/or segment of video,” where, during processing by the second neural network, the rearranged audio derived text sequence is an embedding.; Mahyar, ¶ Col. 4, lines 59-67).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the storyline generation systems of Yun to incorporate the teachings of Mahyar to include generating embeddings of the rearranged audio-derived text sequence. Yun discloses the generation of storylines allowing for novel content generation from existing content, and which further considers discontinuous content. However, Yun is silent as to the automated detection of said discontinuities. Mahyar discloses the automated determination of “non-continuous segments that are related” incorporating neural network models, which allows for automating the determination of the corresponding portions of events, while avoiding the inclusion of incomplete events due to having internal interruptions, which provides the known benefit of reducing costs for manual annotation, and increasing the efficiency of the storyline generation described in Yun, as recognized in light of the disclosure of Mahyar. (Mahyar, ¶ Col. 7, lines 5-16).
Regarding claim 2, the rejection of claim 1 is incorporated. Yun and Mahyar disclose all of the elements of the current invention as stated above. Yun further discloses wherein the segment comprises at least one of a content recap, a content replay, and a content preview (“a storyline 134 following a plot may be combined with other storylines 134” where a storyline is at least a content recap (as defined by the applicant, a recap “present[s] content from a previous episode or segment to explain something(s) that already happened in the narrative or story”) where “the set of one or more tags 132 which define a storyline 134 in the first few episodes of a series” is a content recap, with respect to at least “storyline 134 at the conclusion of the series”; Yun, ¶ col. 4, lines 47-59).
Regarding claim 3, the rejection of claim 1 is incorporated. Yun and Mahyar disclose all of the elements of the current invention as stated above. Yun further discloses wherein the position of the text data within the temporal sequence of the audio-derived text sequence corresponds to the first time in the timeline and the different position of the text data within the audio-derived text sequence corresponds to the second time in the timeline, (Discloses “arranging” the tags, and thus arranging the content, “by increasing time of the chronology of the story” according to the chronological time of the events. As the content is rearranged according to chronological order of the story, the “caption data” corresponding to the flashback is moved from original time position to the chronologically correct time position in the sequence; Yun, ¶ Col. 15, lines 22-36) wherein the text data is moved to the different position within the modified temporal sequence of the rearranged audio-derived text sequence (As described, the content is reorganized by arranging the corresponding tags in chronological order. As the remaining portions of the content are already in the appropriate chronological order, the text data corresponding to the flashback is moved from the first position, which corresponds to the original non-chronological position, to second position, which corresponds to the different chronological position.; Yun, ¶ Col. 15, lines 22-36), and wherein the operations further comprise: … a portion of content from the content item corresponding to a third time in the timeline…[including] data about at least one of a scene, character, dialogue, and event associated with the first time in the timeline (Discloses “Block 1010 generates one or more of tags 132 associated with the content 124 based at least in part on the determined person data and the one or more occurrence times 212” as well as “location data indicative of a location {a scene...} associated with the content 124 at the one or more occurrence times 212 in the content 124 {associated with the first time in the timeline}.”; Yun, ¶ Col. 13, lines 50-67). However, Yun fails to expressly recite wherein the operations further comprise: extracting, from a portion of content from the content item corresponding to a third time in the timeline, data about at least one of a scene, character, dialogue, and event associated with the first time in the timeline; adding the extracted data to a text segment from the rearranged audio-derived text sequence, wherein the text segment corresponds to the first time in the timeline; and generating the embeddings of the rearranged audio-derived text sequence, the rearranged audio-derived text sequence comprising the extracted data added to the text segment.
The relevance of Mahyar is described above with relation to claim 1. Regarding claim 3, Mahyar teaches wherein the operations further comprise: extracting, from a portion of content from the content item corresponding to a third time in the timeline, data about at least one of a scene, character, dialogue, and event associated with the first time in the timeline (Discloses “determin[ing] a first segment of video content, the first segment comprising a first set of frames and first audio content” and “the first segment may be a video segment of video content, and may be associated with text and/or audio components” as part of “non-continuous segments that are related” such as “a scene in the content may be interrupted by a flashback or other scene” which “may correspond to events, scenes, and/or other occurrences that may be discrete and/or extractable from the content” and said segments “may correspond to certain locations and/or times, certain actors that appear, certain music or sounds, and/or other features of the content,” where the system “generate[s], using a second neural network and the vector, a first textual description of the first segment, wherein the first textual description comprises words that describe events of the first segment”; Mahyar, ¶ Col. 6, line 63 - col. 7, line 27); adding the extracted data to a text segment from the rearranged audio-derived text sequence, (“Textual descriptions may be presented... in addition to, subtitles, captions, or other text data.”; Mahyar, ¶ Col. 2, lines 15-32; Col. 5, lines 32-40) wherein the text segment corresponds to the first time in the timeline (The textual description of the first segment can correspond to “a scene in the content” which “may be interrupted by a flashback,” thus, in the context of the flashback presented above from Yun, corresponds to the first time in the timeline.; Mahyar, ¶ Col. 7, lines 5-27); and generating the embeddings of the rearranged audio-derived text sequence (“The second neural network may analyze the individual textual descriptions, aggregate and combine the individual textual descriptions, incorporate analysis of the audio, and output a textual description for the scene and/or segment of video,” where, during processing by the second neural network, the rearranged audio derived text sequence is an embedding.; Mahyar, ¶ Col. 4, lines 59-67), the rearranged audio-derived text sequence comprising the extracted data added to the text segment (Discloses the textual descriptions being added to the “subtitles, captions, or other text data” where the textual descriptions, as understood in the context of Yun, such as at Col. 5, lines 43-46, are understood as content metadata, where “the content metadata 130 may also be used as well to generate the tags 132.”; Mahyar, ¶ Col. 2, lines 15-32; Col. 5, lines 32-40).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the storyline generation systems of Yun to incorporate the teachings of Mahyar to include wherein the operations further comprise: extracting, from a portion of content from the content item corresponding to a third time in the timeline, data about at least one of a scene, character, dialogue, and event associated with the first time in the timeline; adding the extracted data to a text segment from the rearranged audio-derived text sequence, wherein the text segment corresponds to the first time in the timeline; and generating the embeddings of the rearranged audio-derived text sequence, the rearranged audio-derived text sequence comprising the extracted data added to the text segment. Yun discloses the generation of storylines allowing for novel content generation from existing content, and which further considers discontinuous content. However, Yun is silent as to the automated detection of said discontinuities. Mahyar discloses the automated determination of “non-continuous segments that are related” incorporating neural network models, which allows for automating the determination of the corresponding portions of events, while avoiding the inclusion of incomplete events due to having internal interruptions, which provides the known benefit of reducing costs for manual annotation, and increasing the efficiency of the storyline generation described in Yun, as recognized in light of the disclosure of Mahyar. (Mahyar, ¶ Col. 7, lines 5-16).
Regarding claim 4, the rejection of claim 3 is incorporated. Yun and Mahyar disclose all of the elements of the current invention as stated above. However, Yun fails to expressly recite wherein the extracted data comprises contextual data about the at least one of the scene, character, dialogue, and event that is missing from additional text data in the rearranged audio-derived text sequence that corresponds to the first time in the timeline.
The relevance of Mahyar is described above with relation to claim 1. Regarding claim 4, Mahyar teaches wherein the extracted data comprises contextual data about the at least one of the scene, character, dialogue, and event (Discloses “determin[ing] a first segment of video content, the first segment comprising a first set of frames and first audio content” and “the first segment may be a video segment of video content, and may be associated with text and/or audio components” as part of “non-continuous segments that are related” such as “a scene in the content may be interrupted by a flashback or other scene” which “may correspond to events, scenes, and/or other occurrences that may be discrete and/or extractable from the content” and said segments “may correspond to certain locations and/or times, certain actors that appear, certain music or sounds, and/or other features of the content,” where the system “generate[s], using a second neural network and the vector, a first textual description of the first segment, wherein the first textual description comprises words that describe events of the first segment”; Mahyar, ¶ Col. 6, line 63 - col. 7, line 27) that is missing from additional text data in the rearranged audio-derived text sequence that corresponds to the first time in the timeline ( “additional text data” as a portion of the “rearranged audio derived text sequence” is not defined, is not indicated to have any relationship with the text data introduced in claim 1, and can include text portions which are, for example, too small to contain the extracted data. As such, the extracted data is missing from the “additional text data” in at least one embodiment.; Mahyar, ¶ Col. 6, line 63 - col. 7, line 27).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the storyline generation systems of Yun to incorporate the teachings of Mahyar to include wherein the extracted data comprises contextual data about the at least one of the scene, character, dialogue, and event that is missing from additional text data in the rearranged audio-derived text sequence that corresponds to the first time in the timeline. Yun discloses the generation of storylines allowing for novel content generation from existing content, and which further considers discontinuous content. However, Yun is silent as to the automated detection of said discontinuities. Mahyar discloses the automated determination of “non-continuous segments that are related” incorporating neural network models, which allows for automating the determination of the corresponding portions of events, while avoiding the inclusion of incomplete events due to having internal interruptions, which provides the known benefit of reducing costs for manual annotation, and increasing the efficiency of the storyline generation described in Yun, as recognized in light of the disclosure of Mahyar. (Mahyar, ¶ Col. 7, lines 5-16).
Regarding claim 5, the rejection of claim 1 is incorporated. Yun and Mahyar disclose all of the elements of the current invention as stated above. Yun further discloses wherein the rearranged audio-derived text sequence further groups different text data from the audio-derived text sequence within the modified temporal sequence based on one or more relationships between portions of the different text data, (Discloses “arranging” the tags, and thus arranging the content, “by increasing time of the chronology of the story” according to the chronological time of the events. As the content is rearranged according to chronological order of the story, the “caption data” corresponding to the flashback is moved from original time position to the chronologically correct time position in the sequence.; Yun, ¶ Col. 15, lines 22-36) wherein the one or more relationships comprise at least one of a chronological relationship, a contextual relationship, and a common timeline associated with the portions of the different text data and the additional portion of the text data (The text portions are grouped based on chronological order of the story, the relationship being at least a chronological relationship.; Yun, ¶ Col. 15, lines 22-36), and wherein the different portions of the different text data comprise at least one of different caption segments, different transcript segments, and different subtitle segments (“The text data may comprise information from retrieving one or more text captions, recognizing speech with a speech recognition module, or recognizing writing with an optical character recognition module” and may “be based at least in part on closed captions, open captions, subtitles, and so forth”; Yun, ¶ Col. 14, lines 16-25).
Regarding claim 6, the rejection of claim 1 is incorporated. Yun and Mahyar disclose all of the elements of the current invention as stated above. However, Yun fails to expressly recite wherein the rearranged audio-derived text sequence comprises additional text data generated based on image data associated with the content item, wherein the image data comprises at least one of one or more video frames and one or more still images, and wherein the embeddings comprise embeddings of the rearranged audio-derived text sequence including the additional text data generated based on the image data.
The relevance of Mahyar is described above with relation to claim 1. Regarding claim 6, Mahyar teaches wherein the rearranged audio-derived text sequence comprises additional text data generated based on image data associated with the content item, (Discloses “determin[ing] a first segment of video content, the first segment comprising a first set of frames and first audio content” and “the first segment may be a video segment of video content, and may be associated with text and/or audio components” as part of “non-continuous segments that are related” such as “a scene in the content may be interrupted by a flashback or other scene” which “may correspond to events, scenes, and/or other occurrences that may be discrete and/or extractable from the content” and said segments “may correspond to certain locations and/or times, certain actors that appear, certain music or sounds, and/or other features of the content,” where the system “generate[s], using a second neural network and the vector, a first textual description of the first segment, wherein the first textual description comprises words that describe events of the first segment”; Mahyar, ¶ Col. 6, line 63 - col. 7, line 27) wherein the image data comprises at least one of one or more video frames and one or more still images (The “textual description” is generated based on a “video segment of video content {video frames},” and where video frames are one or more still images; Mahyar, ¶ Col. 6, line 63 - col. 7, line 27), and wherein the embeddings comprise embeddings of the rearranged audio-derived text sequence (“The second neural network may analyze the individual textual descriptions, aggregate and combine the individual textual descriptions, incorporate analysis of the audio, and output a textual description for the scene and/or segment of video,” where, during processing by the second neural network, the rearranged audio derived text sequence is an embedding.; Mahyar, ¶ Col. 4, lines 59-67) including the additional text data generated based on the image data (Mahyar discloses the textual descriptions, which are generated based on image data,{ the additional text data generated based on the image data} being added to the “subtitles, captions, or other text data” where the textual descriptions, as understood in the context of Yun, such as at Col. 5, lines 43-46, are understood as content metadata, where “the content metadata 130 may also be used as well to generate the tags 132 {including the additional text data...}.”; Mahyar, ¶ Col. 2, lines 15-32; Col. 5, lines 32-40).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the storyline generation systems of Yun to incorporate the teachings of Mahyar to include wherein the rearranged audio-derived text sequence comprises additional text data generated based on image data associated with the content item, wherein the image data comprises at least one of one or more video frames and one or more still images, and wherein the embeddings comprise embeddings of the rearranged audio-derived text sequence including the additional text data generated based on the image data. Yun discloses the generation of storylines allowing for novel content generation from existing content, and which further considers discontinuous content. However, Yun is silent as to the automated detection of said discontinuities. Mahyar discloses the automated determination of “non-continuous segments that are related” incorporating neural network models, which allows for automating the determination of the corresponding portions of events, while avoiding the inclusion of incomplete events due to having internal interruptions, which provides the known benefit of reducing costs for manual annotation, and increasing the efficiency of the storyline generation described in Yun, as recognized in light of the disclosure of Mahyar. (Mahyar, ¶ Col. 7, lines 5-16).
Regarding claim 7, the rejection of claim 1 is incorporated. Yun and Mahyar disclose all of the elements of the current invention as stated above. Yun further discloses wherein the operations further comprise: detecting the deviation in the timeline of the content item (Discloses “generat[ing] one or more tags 132 associated with the content 124 based at least in part on the determined object data and the one or more occurrence times 212,” where the one or more tags can each indicate a time frame in the content (e.g., a tag for “a saddle appears at times 51:13 to 55:21.”) where the tags are presented as part of a storyline, and the storyline is defined by “associating tags 132 with similar or the same information across time in one or more pieces of the content 124.”; Yun, ¶ Col. 4, lines 41-46; Col. 14, lines 11-15). However, Yun fails to expressly recite based on at least one of the text data corresponding to the segment that corresponds to the deviation in the timeline, the audio associated with the content item, and image data associated with the content item, the image data comprising at least one of one or more video frames and one or more still images.
The relevance of Mahyar is described above with relation to claim 1. Regarding claim 7, Mahyar teaches based on at least one of the text data corresponding to the segment that corresponds to the deviation in the timeline, the audio associated with the content item, and image data associated with the content item, the image data comprising at least one of one or more video frames and one or more still images (“the video processing module(s) 320 may include facial recognition and/or human pose detection algorithms that can be used to identify people or actions in certain locations over frames or segments of the video content” and “through the use of “one or more object recognition algorithms configured to detect at least one of predefined objects, predefined scenery (e.g., certain locations, etc.), and the like,” where the identification may be used to identify “a scene...interrupted by a flashback or cut to a different story”; Mahyar, ¶ Col. 6, line 63 - col. 7, line 27; Col. 10, lines 4-28).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the storyline generation systems of Yun to incorporate the teachings of Mahyar to include based on at least one of the text data corresponding to the segment that corresponds to the deviation in the timeline, the audio associated with the content item, and image data associated with the content item, the image data comprising at least one of one or more video frames and one or more still images. Yun discloses the generation of storylines allowing for novel content generation from existing content, and which further considers discontinuous content. However, Yun is silent as to the automated detection of said discontinuities. Mahyar discloses the automated determination of “non-continuous segments that are related” incorporating neural network models, which allows for automating the determination of the corresponding portions of events, while avoiding the inclusion of incomplete events due to having internal interruptions, which provides the known benefit of reducing costs for manual annotation, and increasing the efficiency of the storyline generation described in Yun, as recognized in light of the disclosure of Mahyar. (Mahyar, ¶ Col. 7, lines 5-16).
Regarding claim 9, the rejection of claim 7 is incorporated. Yun and Mahyar disclose all of the elements of the current invention as stated above. However, Yun fails to expressly recite wherein detecting the deviation in the timeline of the content item comprises: based on the image data, recognizing, using facial recognition, a character depicted in a portion of the image data, wherein the portion of the image data is part of the segment and a respective temporal location of the portion of the image data corresponds to the first time in the timeline; and detecting the deviation in the timeline of the content item based on a determination that the character is associated with the second time in the timeline.
The relevance of Mahyar is described above with relation to claim 1. Regarding claim 9, Mahyar teaches wherein detecting the deviation in the timeline of the content item comprises: based on the image data, recognizing, using facial recognition, a character depicted in a portion of the image data, wherein the portion of the image data is part of the segment (“the video processing module(s) 320 may include facial recognition and/or human pose detection algorithms that can be used to identify people or actions in certain locations over frames or segments of the video content” and may include “a facial recognition module that may be used to analyze video and/or audio of the content in a frame-by-frame or segment-by-segment analysis to detect the presence of characters in frames or scenes” of the video content {image data}; Mahyar, ¶ Col. 6, line 63 - col. 7, line 27; Col. 7, lines 49-67; Col. 10, lines 4-28) and a respective temporal location of the portion of the image data corresponds to the first time in the timeline (“Segments may correspond to certain scenes of the content 410 and… may not be continuous,” also described as “non-continuous segments that are related” such segments including “a flashback,” where said segments “may be identified using one or more timestamps,”; Mahyar, ¶ Col. 13, lines 10-23); and detecting the deviation in the timeline of the content item (the facial identification may be used to identify “a scene...interrupted by a flashback or cut to a different story” through the use of “one or more object recognition algorithms configured to detect at least one of predefined objects, predefined scenery (e.g., certain locations, etc.), and the like.”; Mahyar, ¶ Col. 6, line 63 - col. 7, line 27; Col. 10, lines 4-28) based on a determination that the character is associated with the second time in the timeline (“facial recognition and/or human pose detection algorithms that can be used to identify people or actions in certain locations over frames or segments of the video content, which may not always be consecutive” such as the identification of “a scene may be briefly interrupted by a flashback” where the discontinuity, with respect to the facial recognition and the flashback, indicates that the facial recognition for the “people” as part of the first part of the segment does not correspond with the facial recognition for the character as applied to the flashback, where the person may be detected by facial recognition as related to the first part of the segment, thus corresponding to the first time, or the flashback part of the segment corresponding to the second time. As further exemplified in FIG. 5, “in the first set of frames 510, output 520 of the algorithm may indicate that a man and a woman are facing forward and looking away from each other at a first frame in the first set. In the second frame of the first set, the output may indicate that the man and the woman are closer to each other.”; Mahyar, ¶ Col. 10, lines 4-28; Col. 15, lines 29-47).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the storyline generation systems of Yun to incorporate the teachings of Mahyar to include wherein detecting the deviation in the timeline of the content item comprises: based on the image data, recognizing, using facial recognition, a character depicted in a portion of the image data, wherein the portion of the image data is part of the segment and a respective temporal location of the portion of the image data corresponds to the first time in the timeline; and detecting the deviation in the timeline of the content item based on a determination that the character is associated with the second time in the timeline. Yun discloses the generation of storylines allowing for novel content generation from existing content, and which further considers discontinuous content. However, Yun is silent as to the automated detection of said discontinuities. Mahyar discloses the automated determination of “non-continuous segments that are related” incorporating neural network models, which allows for automating the determination of the corresponding portions of events, while avoiding the inclusion of incomplete events due to having internal interruptions, which provides the known benefit of reducing costs for manual annotation, and increasing the efficiency of the storyline generation described in Yun, as recognized in light of the disclosure of Mahyar. (Mahyar, ¶ Col. 7, lines 5-16).
Regarding claim 10, the rejection of claim 7 is incorporated. Yun and Mahyar disclose all of the elements of the current invention as stated above. However, Yun fails to expressly recite wherein detecting the deviation in the timeline of the content item comprises: based on the image data, recognizing, using scene or image recognition, a scene depicted in a portion of the image data, wherein the portion of the image data is part of the segment and a respective temporal location of the portion of the image data corresponds to the first time in the timeline; and detecting the deviation in the timeline of the content item based on a determination that the scene matches a previous scene in the timeline.
The relevance of Mahyar is described above with relation to claim 1. Regarding claim 10, Mahyar teaches wherein detecting the deviation in the timeline of the content item comprises: based on the image data, recognizing, using scene or image recognition, a scene depicted in a portion of the image data, (“the video processing module(s) 320 may include facial recognition and/or human pose detection algorithms that can be used to identify people or actions in certain locations over frames or segments of the video content, which may not always be consecutive” and may further include “one or more object recognition algorithms configured to detect at least one of predefined objects, predefined scenery” which may be used as part of detecting “certain locations” which may be “used to identify people or actions in certain locations over frames or segments of the video content” where non-consecutive frames can indicate “a flashback or cut to a different story” which interrupts a scene.; Mahyar, ¶ Col. 10, lines 4-28) wherein the portion of the image data is part of the segment and a respective temporal location of the portion of the image data corresponds to the first time in the timeline (The described flashback is part of the segment and the flashback is defined as a flashback, and the temporal time of said flashback corresponds to the first time in the timeline.; Mahyar, ¶ Col. 10, lines 4-28); and detecting the deviation in the timeline of the content item based on a determination that the scene matches a previous scene in the timeline (The system can “determine frames or sets of frames of video content and may be configured to detect certain features, such as certain objects, as well as actions or events across multiple frames” and can “identify people or actions in certain locations” which are non-consecutive as an indication of “a flashback or cut to a different story”, which is understood as including the identification of a first person with relation to a first action in a first frame at some timepoint A, a failure to identify the first person with relation to the first action at some timepoint B, and the later identification of the first person with relation to the first action in a third frame at some timepoint C, and these frames are non-consecutive because of the inclusion of one or more frames at timepoint B which do not maintain the person/action relationship.; Mahyar, ¶ Col. 10, lines 4-28).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the storyline generation systems of Yun to incorporate the teachings of Mahyar to include wherein detecting the deviation in the timeline of the content item comprises: based on the image data, recognizing, using scene or image recognition, a scene depicted in a portion of the image data, wherein the portion of the image data is part of the segment and a respective temporal location of the portion of the image data corresponds to the first time in the timeline; and detecting the deviation in the timeline of the content item based on a determination that the scene matches a previous scene in the timeline. Yun discloses the generation of storylines allowing for novel content generation from existing content, and which further considers discontinuous content. However, Yun is silent as to the automated detection of said discontinuities. Mahyar discloses the automated determination of “non-continuous segments that are related” incorporating neural network models, which allows for automating the determination of the corresponding portions of events, while avoiding the inclusion of incomplete events due to having internal interruptions, which provides the known benefit of reducing costs for manual annotation, and increasing the efficiency of the storyline generation described in Yun, as recognized in light of the disclosure of Mahyar. (Mahyar, ¶ Col. 7, lines 5-16).
Regarding claim 12, Yun discloses A computer-implemented method comprising (Systems and methods described with relation to “generating storylines associated with content” which is implemented using a “device 104” through “one or more processors 106” the processors being “configured to access and execute at least in part instructions stored in the one or more memories 108.”; Yun, ¶ Col. 2, line 64 - col 3, line 33): detecting a deviation in a timeline of a content item and a segment of the content item corresponding to the deviation in the timeline, (Discloses “generat[ing] one or more tags 132 associated with the content 124 based at least in part on the determined object data and the one or more occurrence times 212,” where the one or more tags can each indicate a time frame in the content (e.g., a tag for “a saddle appears at times 51:13 to 55:21.”) where the tags are presented as part of a storyline, and the storyline is defined by “associating tags 132 with similar or the same information across time in one or more pieces of the content 124.”; Yun, ¶ Col. 4, lines 41-46; Col. 14, lines 11-15) wherein a temporal location of the segment corresponds to a first time in the timeline and a content of the segment corresponds to a second time in the timeline (“the tags 132 provide machine-readable information about events 208 taking place at a particular point in the content 124” and “may include reference information, associating the tag with a particular portion of the content 124 or designating the particular portion of the content at which the event described by the tag occurs,” where “each event 208 has at least one corresponding occurrence time which specifies the point or interval in the content at which the event takes place” and a “time” in the “chronology of the story” (hereinafter, a chronological time) where, in the context of a flashback, the tag indicates a temporal location of the segment of the content corresponding to a first time, based on surrounding, non-contiguous events, and the event has a chronological time which is earlier in the story {a content of the segment corresponds to a second time in the timeline}; Yun, ¶ Col. 4, lines 34-40; col. 5, lines 15-23, and 59-67; Col. 15, lines 22-36); identifying text data, from an audio-derived text sequence associated with a content item, that corresponds to the segment of the content item (“determines text data associated with the one or more occurrence times 212 in the content 124,” where the text data is derived from the entire transcript or caption sequence {an audio derived text sequence}; Yun, ¶ Col. 14, lines 16-21), the audio-derived text sequence comprising at least one of a transcript of audio associated with the content item, captions of the audio associated with the content items and subtitles corresponding to the audio associated with the content item (“The text data may comprise information from retrieving one or more text captions, recognizing speech with a speech recognition module, or recognizing writing with an optical character recognition module” and may “be based at least in part on closed captions, open captions, subtitles, and so forth”; Yun, ¶ Col. 14, lines 16-25); based on the second time in the timeline corresponding to the content of the segment and a position of the text data within a temporal sequence of the audio-derived text sequence (Discloses “arranging” the tags “by increasing time of the chronology of the story” according to the chronological time of the events.; Yun, ¶ Col. 15, lines 22-36), determining a different position for the text data within the audio-derived text sequence, (The above arrangement “may result in a storyline which presents the scenes showing the flashbacks first, consistent with their earlier appearance in the internal chronology of the story as being earlier events,” which includes a different determined position for the text data within the storyline corresponding to the chronological occurrence of events within the content, where “The content 124 may comprise one or more of video data 202, audio data 204, [and] caption data 206 {audio derived text sequence}”; Yun, ¶ Col. 4, lines 60-63; Col. 15, lines 22-36) wherein the different position is associated with the second time in the timeline (The different position is “consistent with their earlier appearance in the internal chronology of the story as being earlier events,” where the chronological time is the second time in the “internal chronology {timeline}” where, in the context of “video data 202, audio data 204, [and] caption data 206” as the content, a different position is determined for each of these, and said position is associated with the second time.; Yun, ¶ Col. 4, lines 60-63; Col. 15, lines 22-36); determining a modified temporal sequence of the audio-derived text sequence (As above, the “storyline... presents the scenes showing the flashbacks first, consistent with their earlier appearance in the internal chronology of the story as being earlier events,” which includes a modified temporal sequence of the events (which necessarily includes a modified temporal sequence of the text sequence in the content, for each of “video data 202, audio data 204, [and] caption data 206”; Yun, ¶ Col. 4, lines 60-63; Col. 15, lines 22-36) based on the temporal sequence of the audio-derived text sequence and the different position within the audio-derived text sequence determined for the text data (Said modified temporal sequence is determined based on the content, where content can be any or all of “video data 202, audio data 204, [and] caption data 206”. As such, the temporal sequence can be based on the “internal chronology” of video, audio, caption (text), or any combination.; Yun, ¶ Col. 4, lines 60-63; Col. 15, lines 22-36); based on the audio-derived text sequence and the modified temporal sequence, generating a rearranged audio-derived text sequence associated with the content item (The “scenes” are presented {generated} “showing the flashbacks first, consistent with their earlier appearance in the internal chronology of the story as being earlier events,” as understood in the context of the caption data 206, this also includes the generation of the caption data rearranged “consistent with their earlier appearance in the internal chronology of the story”; Yun, ¶ Col. 15, lines 22-36), the rearranged audio-derived text sequence comprising the audio-derived text sequence rearranged according to the modified temporal sequence (“storyline 134 may be configured to present portions of the content designated by the plurality of tags 132 while omitting portions of the content 124 undesignated by the plurality of tags 132” which is understood to include not “omitting portions of the content 124 undesignated by the plurality of tags 132” and “a storyline 134 for a character which features may flashback and flash-forward scenes may result in a storyline which presents the scenes showing the flashbacks first, consistent with their earlier appearance in the internal chronology of the story” which is a rearrangement of the content {audio derived text sequence} according to the order of the event’s chronological appearance {the modified temporal sequence}; Yun, ¶ Col. 15, lines 22-36). However, Yun fails to expressly recite generating embeddings of the rearranged audio-derived text sequence.
The relevance of Mahyar is described above with relation to claim 1. Regarding claim 12, Mahyar teaches generating embeddings of the rearranged audio-derived text sequence (“The second neural network may analyze the individual textual descriptions, aggregate and combine the individual textual descriptions, incorporate analysis of the audio, and output a textual description for the scene and/or segment of video,” where, during processing by the second neural network, the rearranged audio derived text sequence is an embedding.; Mahyar, ¶ Col. 4, lines 59-67).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the storyline generation systems of Yun to incorporate the teachings of Mahyar to include generating embeddings of the rearranged audio-derived text sequence. Yun discloses the generation of storylines allowing for novel content generation from existing content, and which further considers discontinuous content. However, Yun is silent as to the automated detection of said discontinuities. Mahyar discloses the automated determination of “non-continuous segments that are related” incorporating neural network models, which allows for automating the determination of the corresponding portions of events, while avoiding the inclusion of incomplete events due to having internal interruptions, which provides the known benefit of reducing costs for manual annotation, and increasing the efficiency of the storyline generation described in Yun, as recognized in light of the disclosure of Mahyar. (Mahyar, ¶ Col. 7, lines 5-16).
Regarding claim 13, the rejection of claim 12 is incorporated. Yun and Mahyar disclose all of the elements of the current invention as stated above. Yun further discloses wherein the rearranged audio-derived text sequence further groups different text data from the audio-derived text sequence within the modified temporal sequence (The “selection of the one or more tags 132 may be based on one or more criteria applied to the description 210 and in some implementations the occurrence time 212” including “sorting the tags 132 by the occurrence time 212 in ascending order from lowest occurrence time 212” where order of occurrence time includes addressing “flashback and flash-forward scenes may result in a storyline which presents the scenes” such as by “showing the flashbacks first, consistent with their earlier appearance in the internal chronology of the story as being earlier events” where the order of appearance is an association of the deviation with additional tags; Yun, ¶ Col. 15, lines 15-36) based on one or more relationships between portions of the different text data, (“In implementations where the correlation between the tags 132 is to be considered, the criteria for the selecting the one or more tags may include the tags 132 having a determined correlation with one another greater than the correlation threshold 704,” where the determined correlation is the one or more relationships.; Yun, ¶ Col. 15, lines 37-41) wherein the one or more relationships comprise at least one of a topic, a chronological relationship, a contextual relationship, and a common timeline (In addition to following the “internal chronology of the story”, the “storyline 134 may combine tags 132 which are associated with different events 208” based on a determined “correlation between two or more tags 132” for example, “the tag 132(1) for the location ‘Ed’s Mercantile’ may have a correlation with the tag 132(3) for the character ‘Ed’” and/or “appearance of “Chet” and “Johnny” may be highly correlated, and a storyline 134 may be generated as described next which includes tags for both characters,” and/or the a storyline of “riding may be based on a combination of tags indicating cowboys and horses,” and tags corresponding to the “events 208 may include...dialogue or discussion on a particular topic”; Yun, ¶ Col. 5, lines 7-14; Col. 14, line 56-Col. 15, line 9).
Regarding claim 14, the rejection of claim 12 is incorporated. Yun and Mahyar disclose all of the elements of the current invention as stated above. Yun further discloses wherein the segment comprises at least one of a content recap, a content replay, a flashback, a flashforward, and a content preview (“a storyline 134 for a character which features [a] flashback” where the flashback is the deviation in the playback timeline; Yun, ¶ Col. 15, lines 22-36)
Regarding claim 15, the rejection of claim 12 is incorporated. Claim 15 is substantially the same as claim 6 and is therefore rejected under the same rationale as above.
Regarding claim 16, the rejection of claim 12 is incorporated. Claim 16 is substantially the same as claim 7 and is therefore rejected under the same rationale as above.
Regarding claim 17, the rejection of claim 16 is incorporated. Claim 17 is substantially the same as claim 9 and is therefore rejected under the same rationale as above.
Regarding claim 18, the rejection of claim 16 is incorporated. Claim 18 is substantially the same as claim 10 and is therefore rejected under the same rationale as above.
Regarding claim 20, Yun discloses A non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform operations comprising (Systems and methods described with relation to “ generating storylines associated with content” which is implemented using a “device 104” through “one or more processors 106” the processors being “configured to access and execute at least in part instructions stored in the one or more memories 108.”; Yun, ¶ Col. 2, line 64 - col 3, line 33): detecting a deviation in a timeline of a content item and a segment of the content item corresponding to the deviation in the timeline, (Discloses “generat[ing] one or more tags 132 associated with the content 124 based at least in part on the determined object data and the one or more occurrence times 212,” where the one or more tags can each indicate a time frame in the content (e.g., a tag for “a saddle appears at times 51:13 to 55:21.”) where the tags are presented as part of a storyline, and the storyline is defined by “associating tags 132 with similar or the same information across time in one or more pieces of the content 124.”; Yun, ¶ Col. 4, lines 41-46; Col. 14, lines 11-15) wherein a temporal location of the segment corresponds to a first time in the timeline and a content of the segment corresponds to a second time in the timeline (“the tags 132 provide machine-readable information about events 208 taking place at a particular point in the content 124” and “may include reference information, associating the tag with a particular portion of the content 124 or designating the particular portion of the content at which the event described by the tag occurs,” where “each event 208 has at least one corresponding occurrence time which specifies the point or interval in the content at which the event takes place” and a “time” in the “chronology of the story” (hereinafter, a chronological time) where, in the context of a flashback, the tag indicates a temporal location of the segment of the content corresponding to a first time, based on surrounding, non-contiguous events, and the event has a chronological time which is earlier in the story {a content of the segment corresponds to a second time in the timeline}; Yun, ¶ Col. 4, lines 34-40; col. 5, lines 15-23, and 59-67; Col. 15, lines 22-36); identifying text data, from an audio-derived text sequence associated with a content item, that corresponds to the segment of the content item (“determines text data associated with the one or more occurrence times 212 in the content 124,” where the text data is derived from the entire transcript or caption sequence {an audio derived text sequence}; Yun, ¶ Col. 14, lines 16-21), the audio-derived text sequence comprising at least one of a transcript of audio associated with the content item, captions of the audio associated with the content items and subtitles corresponding to the audio associated with the content item (“The text data may comprise information from retrieving one or more text captions, recognizing speech with a speech recognition module, or recognizing writing with an optical character recognition module” and may “be based at least in part on closed captions, open captions, subtitles, and so forth”; Yun, ¶ Col. 14, lines 16-25); based on the second time in the timeline corresponding to the content of the segment and a position of the text data within a temporal sequence of the audio-derived text sequence (Discloses “arranging” the tags “by increasing time of the chronology of the story” according to the chronological time of the events.; Yun, ¶ Col. 15, lines 22-36), determining a different position for the text data within the audio-derived text sequence, (The above arrangement “may result in a storyline which presents the scenes showing the flashbacks first, consistent with their earlier appearance in the internal chronology of the story as being earlier events,” which includes a different determined position for the text data within the storyline corresponding to the chronological occurrence of events within the content, where “The content 124 may comprise one or more of video data 202, audio data 204, [and] caption data 206 {audio derived text sequence}”; Yun, ¶ Col. 4, lines 60-63; Col. 15, lines 22-36) wherein the different position is associated with the second time in the timeline (The different position is “consistent with their earlier appearance in the internal chronology of the story as being earlier events,” where the chronological time is the second time in the “internal chronology {timeline}” where, in the context of “video data 202, audio data 204, [and] caption data 206” as the content, a different position is determined for each of these, and said position is associated with the second time.; Yun, ¶ Col. 4, lines 60-63; Col. 15, lines 22-36); determining a modified temporal sequence of the audio-derived text sequence (As above, the “storyline... presents the scenes showing the flashbacks first, consistent with their earlier appearance in the internal chronology of the story as being earlier events,” which includes a modified temporal sequence of the events (which necessarily includes a modified temporal sequence of the text sequence in the content, for each of “video data 202, audio data 204, [and] caption data 206”; Yun, ¶ Col. 4, lines 60-63; Col. 15, lines 22-36) based on the temporal sequence of the audio-derived text sequence and the different position within the audio-derived text sequence determined for the text data (Said modified temporal sequence is determined based on the content, where content can be any or all of “video data 202, audio data 204, [and] caption data 206”. As such, the temporal sequence can be based on the “internal chronology” of video, audio, caption (text), or any combination.; Yun, ¶ Col. 4, lines 60-63; Col. 15, lines 22-36); based on the audio-derived text sequence and the modified temporal sequence, generating a rearranged audio-derived text sequence associated with the content item (The “scenes” are presented {generated} “showing the flashbacks first, consistent with their earlier appearance in the internal chronology of the story as being earlier events,” as understood in the context of the caption data 206, this also includes the generation of the caption data rearranged “consistent with their earlier appearance in the internal chronology of the story”; Yun, ¶ Col. 15, lines 22-36), the rearranged audio-derived text sequence comprising the audio-derived text sequence rearranged according to the modified temporal sequence (“storyline 134 may be configured to present portions of the content designated by the plurality of tags 132 while omitting portions of the content 124 undesignated by the plurality of tags 132” which is understood to include not “omitting portions of the content 124 undesignated by the plurality of tags 132” and “a storyline 134 for a character which features may flashback and flash-forward scenes may result in a storyline which presents the scenes showing the flashbacks first, consistent with their earlier appearance in the internal chronology of the story” which is a rearrangement of the content {audio derived text sequence} according to the order of the event’s chronological appearance {the modified temporal sequence}; Yun, ¶ Col. 15, lines 22-36). However, Yun fails to expressly recite generating embeddings of the rearranged audio-derived text sequence.
The relevance of Mahyar is described above with relation to claim 1. Regarding claim 20, Mahyar teaches generating embeddings of the rearranged audio-derived text sequence (“The second neural network may analyze the individual textual descriptions, aggregate and combine the individual textual descriptions, incorporate analysis of the audio, and output a textual description for the scene and/or segment of video,” where, during processing by the second neural network, the rearranged audio derived text sequence is an embedding.; Mahyar, ¶ Col. 4, lines 59-67).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the storyline generation systems of Yun to incorporate the teachings of Mahyar to include generating embeddings of the rearranged audio-derived text sequence. Yun discloses the generation of storylines allowing for novel content generation from existing content, and which further considers discontinuous content. However, Yun is silent as to the automated detection of said discontinuities. Mahyar discloses the automated determination of “non-continuous segments that are related” incorporating neural network models, which allows for automating the determination of the corresponding portions of events, while avoiding the inclusion of incomplete events due to having internal interruptions, which provides the known benefit of reducing costs for manual annotation, and increasing the efficiency of the storyline generation described in Yun, as recognized in light of the disclosure of Mahyar. (Mahyar, ¶ Col. 7, lines 5-16).
Regarding claim 21, the rejection of claim 12 is incorporated. Claim 21 is substantially the same as claim 3 and is therefore rejected under the same rationale as above.
Claims 11 and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yun and Mahyar as applied to claim 7 and 16 above, and further in view of Non-patent Literature to Eisenberg (Eisenberg, J & Finlayson, M. (2021) 'Narrative Boundaries Annotation Guide', Journal of Cultural Analytics. 6(4) https://doi.org/10.22148/001c.30698, hereinafter Eisenberg).
Regarding claim 11, the rejection of claim 7 is incorporated. Yun and Mahyar disclose all of the elements of the current invention as stated above. Yun further discloses wherein detecting the deviation in the timeline of the content item comprises: recognizing, using speech or voice recognition, at least one of an utterance in the audio associated with the content item, speech in the audio associated with the content item, and a voice in the audio associated with the content item (Discloses “A recognition module 624... configured to recognize events 208 or event constituents within the content 124” and which “generates data about components of the events 208” using “facial recognition, voice recognition to identify a particular speaker, speech recognition to transform spoken words to text for processing, object recognition, optical character recognition, and so forth.; Yun, ¶ Col. 9, lines 45-59). However, Yun and Mahyar fail to expressly recite detecting the deviation in the timeline of the content item based on a determination that at least one of the voice and a character associated with the voice is associated with the second time in the timeline and a portion of the audio containing the voice corresponds to the first time in the timeline.
Eisenberg teaches systems and methods for training computers “to identify the beginnings and ends of narratives”. (Eisenberg, ¶ Pg. 1, para. 1). Regarding claim 11, Eisenberg teaches detecting the deviation in the timeline of the content item based on a determination that at least one of the voice and a character associated with the voice is associated with the second time in the timeline and a portion of the audio containing the voice corresponds to the first time in the timeline (Discloses the detection of interrupt narratives, which “can occur within chapters, or, for our purposes, within short stories, chapters of novels, or in the dialogue of a script” where a change in “the person narrating” indicates “an interruptive narrative boundary.” Where the determination of the person narrating, in the context of the “dialogue of a script” is the at least one voice and the character associated with the voice, each of which are associated with “the story of the original narrative {associated with a second time in the timeline}” and this occurs before “the interrupting narrative begins to be told” as the interrupting narrative begins “Once the original narrative has stopped {audio containing the voice corresponds to the first time in the timeline}”; Eisenberg, ¶ pg. 36-37, inclusive).
It would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the storyline generation systems of Yun, as modified by the automated video discontinuity detection systems of Mahyar to incorporate the teachings of Eisenberg to include detecting the deviation in the timeline of the content item based on a determination that at least one of the voice and a character associated with the voice is associated with the second time in the timeline and a portion of the audio containing the voice corresponds to the first time in the timeline. Yun discloses the use of “voice recognition to identify a particular speaker” within the content, and the recognition of events including flashbacks within the content. However, Yun is silent as to the automated detection of said discontinuities. Eisenberg discloses using the collected audio information for the detection of interrupt narratives, which includes interrupt flashbacks and flashforwards. It would have been obvious to one having ordinary skill in the art to combine the known techniques from Yun for automatically detecting a particular speaker and determining the content of said speech using the disclosed speech recognition, to the detection of interrupt narratives based on a speaker (voice, speech, utterance) being different between two scenes, as the further use of audio and text to determine sequence for the tags based on chronological order further improves “the user experience... allowing the user to consume content which is most meaningful to them,” by automating the detection of narrative boundaries as disclosed in Eisenberg, not least of which by increasing the amount of available content to which the techniques of Yun may be applied, as recognized in light of the disclosure of Eisenberg. (Eisenberg, ¶ pg. 1, Abstract and Introduction).
Regarding claim 19, the rejection of claim 16 is incorporated. Claim 19 is substantially the same as claim 11 and is therefore rejected under the same rationale as above.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Thomas et al. (U.S. Pat. App. Pub. No. 2021/0385546) discloses systems and methods for allowing users to use a media guidance application to view and navigate customized media presentations.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Sean E. Serraguard whose telephone number is (313)446-6627. The examiner can normally be reached 07:00-17:00 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel C. Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Sean E Serraguard/Primary Examiner, Art Unit 2657