DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1,148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating
obviousness or nonobviousness.
4. Claims 1-2, 8, 11, and 14-20 are rejected under AIA 35 U.S.C. 103 as being unpatentable over Fu (English Translation of Chinese Publication CN111182358 05-2000) in view of Anderton-Yang (US Publication 2022/0237892), and further in view of Brauckmann et al. (US Publication 2017/0294213).
Regarding claim 1, Fu discloses a method, comprising:
recording, by an application server, a multimedia content of a live event (Fu, para. 0010, recording video of a live event);
determining, by the application server, a first alert associated with the multimedia content at a first time instance, wherein the first alert is indicative of a start of a portion of the multimedia content that is to be tagged; determining, by the application server, a second time instance that corresponds to an end of the portion of the multimedia content that is to be tagged; generating, by the application server, a first multimedia clip based on the multimedia content, wherein the first multimedia clip includes the portion of the multimedia content that is to be tagged (Fu, para. 0074-0076, after obtaining the recorded video, performing content recognition on the recorded video; obtaining/generating n video segments based on the content recognition results. The above content recognition is used to identify the content included in the recorded video, which may include songs, jokes, talent performances, audience comments, etc. “determining an alert”; para. 0093, the above-mentioned content recognition of the recorded video to obtain n video segments may include: identifying, from the recorded video, the start and end times, i.e., “a first time instance and a second time instance”, of a song, applause for a joke, a message, a live chat, gift-giving, cheers, etc., extracting the video segment between the start and end times of the song, applause for a joke, a message, a live chat, gift-giving, cheers, etc., and obtaining video segment of a song, a joke, a message, a live chat, gift giving, victory, etc.; para’s 0097-0101, after obtaining the above n video segments, the tag information of each video segment can be obtained; as such, the obtained/generated video segment includes the song, the joke, the message, the live chat, the gift giving, or the victory in the multimedia content that is to be tagged);
identifying, by the application server, from a plurality of tags, a first tag that is indicative of a context of the first multimedia clip (Fu, para’s 00097-0101, Step 303: after obtaining the above n video segments, the tag information of each video segment can be obtained. The above tag information is used to characterize the features of video segments including content tags, which are used to indicate the content category corresponding to the video clip. For example, the content tag may include at least one of the following: song, joke, comment, gift, competition victory, talent. Step 304: determine weight corresponding to each of the n video segments based on the label information of the n video segments. The determined weights are used to indicate the priority of the video segments; para’s 0105-0133, for each of the i-th video segment of the n video segments, k tags can be obtained “the plurality of tags”, and the weight score corresponding to each tag can be obtained. Determine the selection weight corresponding to the i-th video segment based on the weight scores corresponding to the k tags respectively. After obtaining the weight scores corresponding to each label/tag, the selection weight corresponding to the i-th video segment can be obtained based on the weight scores corresponding to each label/tag).
Fu does not explicitly disclose:
the multimedia content that is indicative of a live user interview of a user;
linking, by the application server, the first tag with the first multimedia clip; and storing, by the application server, the first multimedia clip and the corresponding first tag in a memory associated with the application server.
Anderton-Yang discloses the multimedia content that is indicative of a user interview of a user (Anderton-Yang, para’s 0071, 0088, and 0092, the content is an interview of a user).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Anderton-Yang’s features into Fu’s invention for enhancing user’s viewing experience by diversifying the scope of the invention to different areas.
Fu-Anderton-Yang does not explicitly disclose but Brauckmann discloses linking, by the application server, the first tag with the first multimedia clip; and storing, by the application server, the first multimedia clip and the corresponding first tag in a memory associated with the application server (Brauckmann, para’s 0007-0010, storing the tag in association with the video data, in particular in a header of a file comprising the video data or including a link between the tag and the video data. A tag is a piece of information that is related to the video data, in particular information related to an object of interest therein, such a keyword or a short text/description).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Brauckmann’s features into Fu-Anderton-Yang’s invention for enhancing user’s video searching experience by providing video index of segments of interest.
Regarding claim 2, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1, wherein the first alert is determined based on detection of a trigger in the live user interview, and wherein the trigger includes at least one of a gesture, a facial expression, and one or more predefined keywords associated with the user in the live user interview (Fu, para. 0093, gesture can be applauses).
Regarding claim 8, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1, further comprising: identifying, by the application server, from the plurality of tags, a second tag that is indicative of the context of the first multimedia clip; linking, by the application server, the second tag with the first multimedia clip; and storing, by the application server, the first multimedia clip and the corresponding second tag in the memory associated with the application server (Fu, para. 0093, the above-mentioned content recognition of the recorded video to obtain n video segments may include: identifying, from the recorded video, the start and end times, i.e., “a first time instance and a second time instance”, of a song, applause for a joke, a message, a live chat, gift-giving, cheers, etc., extracting the video segment between the start and end times of the song, applause for a joke, a message, a live chat, gift-giving, cheers, etc., and obtaining video segment of a song, a joke, a message, a live chat, gift giving, victory, etc.; para’s 0097-0101, after obtaining the above n video segments, the tag information of each video segment can be obtained; the above tag information is used to characterize the features of video segments including content tags, which are used to indicate the content category corresponding to the video clip. For example, the content tag may include at least one of the following: song, joke, comment, gift, competition victory, talent; as such, the obtained/generated video segment includes video context, for example, the song, the joke, the message, the live chat, the gift giving, or the victory in the multimedia content that is to be tagged; as such, a second content tag associated with different context of the video clip may be obtained; Brauckmann, para’s 0007-0010, storing the tag in association with the video data, in particular in a header of a file comprising the video data or including a link between the tag and the video data. A tag is a piece of information that is related to the video data, in particular information related to an object of interest therein, such as a keyword or a short text/description; as such, a second content tag associated with different object(s) in the video clip may be obtained).
The motivation to combine the references and obviousness arguments are the same as claim 1.
Regarding claim 11, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1, wherein the user can be an organizer of a live user interview (see Anderton-Yang, para. 0078), the method further comprises: receiving, by the application server, after the first alert, a label, wherein the label corresponds to one or more characteristics, that are different from the first tag, assigned to the portion of the multimedia content that is to be tagged; linking, by the application server, the label with the first multimedia clip; and storing, by the application server, in conjunction with the first multimedia clip and the first tag, the label in the memory associated with the application server (Fu, para. 0093, the above-mentioned content recognition of the recorded video to obtain n video segments may include: identifying, from the recorded video, the start and end times, i.e., “a first time instance and a second time instance”, of a song, applause for a joke, a message, a live chat, gift-giving, cheers, etc., extracting the video segment between the start and end times of the song, applause for a joke, a message, a live chat, gift-giving, cheers, etc., and obtaining video segment of a song, a joke, a message, a live chat, gift giving, victory, etc.; para’s 0097-0101, after obtaining the above n video segments, the tag information of each video segment can be obtained; the above tag information is used to characterize the features of video segments including content tags, which are used to indicate the content category corresponding to the video clip. For example, the content tag may include at least one of the following: song, joke, comment, gift, competition victory, talent; as such, the obtained/generated video segment includes video context, for example, the song, the joke, the message, the live chat, the gift giving, or the victory in the multimedia content that is to be tagged; para’s 0100-0107, identifying/determining label for each segment; as such, after a first alert, a label that corresponds to one or more characteristics, that are different from the first tag, assigned to the portion of the multimedia content that is to be tagged may be obtained; Anderton-Yang, para’s 0062 and 0066, generating user-defined labels and system-generated labels for the recorded vide. Labels for the videos may be changed. As an example, the system can obtain information from social media profiles of users who previously recorded practice interview videos and determine from the information whether the users were successful in actual interviews. If a user is determined to have succeeded in an actual interview, then the user's videos that are inferred to be similar to the successful video (e.g., due to ratings of the user or being created at around a similar time) can be updated to indicate that the videos should be classified in a performance classification for high quality or high likelihood of a successful outcome. The machine learning model can then be retrained using the updated training data that includes these updated labels. The target feature values, ranges, or series can also be updated using the feature values extracted from videos of users who were determined to be successful in their actual interviews; Brauckmann, para’s 0007-0010, storing the tag or the label in association with the video data, in particular in a header of a file comprising the video data or including a link between the tag or the label and the video data. As known in the art, a tag and a label individually is a piece of information that is related to the video data, in particular information related to an object of interest therein, such as a keyword or a short text/description; as such, a label can be stored in conjunction with a video and a tag that corresponds to the video content).
The motivation to combine the references and well-known technique in the art and obviousness arguments are the same as claim 1.
Regarding claim 14, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1, wherein the user can be an organizer of a live user interview (see Anderton-Yang, para. 0078), the method further comprises: presenting, by the application server, on an organizer device of an organizer of the live user interview, the multimedia content having the first multimedia clip and the first tag; and receiving, by the application server, via the organizer device, an input that verifies the first tag and the start and the end of the first multimedia clip (Fu, para’s 0146, 0166-0168, display video and other information of the video; Anderton-Yang, para’s 0092, 0177, displaying video content; Brauckmann, para. 0012, a plurality of video data may be provided by a plurality of video recording devices and the step of displaying may comprise displaying a map of the location of the video recording devices and/or displaying a time bar indicating availability of video data from the plurality of video recording devices at different times and/or for a particular view area of interest; receiving input or instruction for verifying the tag and the start and the end of the media clip is well known in the art).
The motivation to combine the references and well-known technique in the art and obviousness arguments are the same as claim 1.
Regarding claim 15, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1, further comprising, receiving, by the application server, over a communication network, the multimedia content from a user device of the user during the live user interview, wherein the live user interview is conducted for gathering user information regarding a domain of the live user interview from the user, and wherein the received multimedia content is recorded to enable the tagging of the multimedia content (Fu, para’s 0072, 0089, 0212, communication network for receiving and transmitting video content; para’s 00097-0101, Step 303: after obtaining the above n video segments, the tag information of each video segment can be obtained; Anderton-Yang, para. 0066, the machine learning model is updated over time using updated training data. The updated training data can include training data generated using different or additional videos, such as those newly recorded or recorded by new users. Alternatively, the updated training data can include training data generated from the same set of videos that were previously used to train the machine learning model. For example, the updated training data can include the same sets of extracted feature values that were previously used to train the machine learning model “domain of the interview”, however the labels for the videos may be changed. As an example, the system can obtain information from social media profiles of users who previously recorded practice interview videos and determine from the information whether the users were successful in actual interviews. If a user is determined to have succeeded in an actual interview, then the user's videos that are inferred to be similar to the successful video (e.g., due to ratings of the user or being created at around a similar time) can be updated to indicate that the videos should be classified in a performance classification for high quality or high likelihood of a successful outcome. The machine learning model can then be retrained using the updated training data that includes these updated labels. The target feature values, ranges, or series can also be updated using the feature values extracted from videos of users who were determined to be successful in their actual interviews; see also para. 0073, the server system 120 can include one or more computing devices, such as one or more servers. The server system 120 can communicate with the computing device 104 over the network 140. The server system 120 can communicate with other computing devices, such as those belonging to other users who have previously recorded practice interview videos or those who intend to record a practice interview video).
The motivation to combine the references and obviousness arguments are the same as claim 1.
Regarding claim 16, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1, wherein the user can be an organizer of a live user interview (see Anderton-Yang, para. 0078), the method further comprises, receiving, over a communication network, the multimedia content from an organizer device of an organizer of the live user interview, wherein the live user interview is conducted for gathering user information regarding a domain of the live user interview from the user, and wherein the received multimedia content is recorded to enable the tagging of the multimedia content (Fu, para’s 0072, 0089, 0212, communication network for receiving and transmitting video content; para’s 00097-0101, Step 303: after obtaining the above n video segments, the tag information of each video segment can be obtained; Anderton-Yang, para. 0066, the machine learning model is updated over time using updated training data. The updated training data can include training data generated using different or additional videos, such as those newly recorded or recorded by new users. Alternatively, the updated training data can include training data generated from the same set of videos that were previously used to train the machine learning model. For example, the updated training data can include the same sets of extracted feature values that were previously used to train the machine learning model “domain of the interview”, however the labels for the videos may be changed. As an example, the system can obtain information from social media profiles of users who previously recorded practice interview videos and determine from the information whether the users were successful in actual interviews. If a user is determined to have succeeded in an actual interview, then the user's videos that are inferred to be similar to the successful video (e.g., due to ratings of the user or being created at around a similar time) can be updated to indicate that the videos should be classified in a performance classification for high quality or high likelihood of a successful outcome. The machine learning model can then be retrained using the updated training data that includes these updated labels. The target feature values, ranges, or series can also be updated using the feature values extracted from videos of users who were determined to be successful in their actual interviews; para. 0073, the server system 120 can include one or more computing devices, such as one or more servers. The server system 120 can communicate with the computing device 104 over the network 140. The server system 120 can communicate with other computing devices, such as those belonging to other users who have previously recorded practice interview videos or those who intend to record a practice interview video).
The motivation to combine the references and obviousness arguments are the same as claim 1.
Regarding claim 17, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1, wherein the user can be an organizer of a live user interview (see Anderton-Yang, para. 0078), the method further comprises, receiving, by the application server, via an organizer device of an organizer of the live user interview, a domain of the live user interview, and the plurality of tags associated with the domain, wherein each tag of the plurality of tags is indicative of at least one of a subject, an objective, and a keyword associated with the domain of the live user interview (Fu, para’s 0072, 0089, 0212, communication network for receiving and transmitting video content; para’s 00097-0101, Step 303: after obtaining the above n video segments, the tag information of each video segment can be obtained; para’s 00097-0101, Step 303: after obtaining the above n video segments, the tag information of each video segment can be obtained. The above tag information is used to characterize the features of video segments including content tags, which are used to indicate the content category corresponding to the video clip. For example, the content tag may include at least one of the following: song, joke, comment, gift, competition victory, talent.; Anderton-Yang, para. 0066, the machine learning model is updated over time using updated training data. The updated training data can include training data generated using different or additional videos, such as those newly recorded or recorded by new users. Alternatively, the updated training data can include training data generated from the same set of videos that were previously used to train the machine learning model. For example, the updated training data can include the same sets of extracted feature values that were previously used to train the machine learning model “domain of the interview”, however the labels for the videos may be changed. As an example, the system can obtain information from social media profiles of users who previously recorded practice interview videos and determine from the information whether the users were successful in actual interviews. If a user is determined to have succeeded in an actual interview, then the user's videos that are inferred to be similar to the successful video (e.g., due to ratings of the user or being created at around a similar time) can be updated to indicate that the videos should be classified in a performance classification for high quality or high likelihood of a successful outcome. The machine learning model can then be retrained using the updated training data that includes these updated labels. The target feature values, ranges, or series can also be updated using the feature values extracted from videos of users who were determined to be successful in their actual interviews; para. 0073, the server system 120 can include one or more computing devices, such as one or more servers. The server system 120 can communicate with the computing device 104 over the network 140. The server system 120 can communicate with other computing devices, such as those belonging to other users who have previously recorded practice interview videos or those who intend to record a practice interview video; para’s 0086-0078, detecting keywords and objects in the video).
The motivation to combine the references and obviousness arguments are the same as claim 1.
Regarding claim 18, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1, further comprising: generating, by the application server, a second multimedia clip that includes a portion of the multimedia content corresponding to a time interval between a third time instance and a fourth time instance; identifying, by the application server, from the plurality of tags, a second tag that is indicative of a context of the second multimedia clip; linking, by the application server, the second tag with the second multimedia clip; determining, by the application server, one or more insights associated with the multimedia content based on an analysis of (i) the first multimedia clip and the corresponding first tag and (ii) the second multimedia clip and the corresponding second tag; and presenting, by the application server, on an organizer device of an organizer of the live user interview, the one or more insights to the organizer (Fu, para’s 0131, 0166, 0191, a third and fourth time instances for a second different segment can be set; para’s 00097-0101, after obtaining the above video segment, the tag information of each video segment can be obtained. The above tag information is used to characterize the features of video segments including content tags, which are used to indicate the content category corresponding to the video clip. For example, the content tag may include at least one of the following: song, joke, comment, gift, competition victory, talent. a second tag associated with the third and the fourth time instances can be obtained; para’s 0105-0133, for each of the i-th video segment of the n video segments, k tags can be obtained “the plurality of tags”, and the weight score corresponding to each tag can be obtained. Determine the selection weight corresponding to the i-th video segment based on the weight scores corresponding to the k tags respectively. After obtaining the weight scores corresponding to each label/tag, the selection weight corresponding to the i-th video segment can be obtained based on the weight scores corresponding to each label/tag; step 305: Select m video segments from n video segments in descending order of selection weight. Specifically, after determining the selection weight corresponding to each of the i-th video segment of the n video segments, m video segments can be selected from the n video segments in descending order of selection weight. For example, the video segments with the highest weights can be selected based on analysis of the n video segment with respective content tag; Brauckmann, para’s 0007-0010, storing the tag in association with the video data, in particular in a header of a file comprising the video data or including a link between the tag and the video data. A tag is a piece of information that is related to the video data, in particular information related to an object of interest therein, such a keyword or a short text/description).
The motivation to combine the references and obviousness arguments are the same as claim 1.
Claims 19-20 are rejected for the same reasons as discussed in claim 1; Fu-Anderton-Yang-Brauckmann further discloses processor(s), memory, and computer readable medium (see Fu, para. 0204).
5. Claims 3-7 are rejected under AIA 35 U.S.C. 103 as being unpatentable over Fu-Anderton-Yang-Brauckmann, as applied to claim 1 above, in view of Khalil et al. (US Publication2024/0292073).
Regarding claim 3, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1, wherein the user can be an organizer of a live user interview (see Anderton-Yang, para. 0078).
Fu-Anderton-Yang-Brauckmann does not explicitly disclose but Khalil discloses wherein the first alert corresponds to an input received via an organizer device of an organizer of the live user interview while the multimedia content is being recorded (Khalil, para. 0071, the user may provide input to enable or disable the variable length, and may also provide input during recording indicating the number of frames or length of time that may be defined as variable start or end frames at the start and/or end of the customized video segment).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Khalil’s features into Fu-Anderton-Yang-Brauckmann’s invention for enhancing user’s video searching experience by allowing user to indicate start point of segments to be tagged.
Regarding claim 4, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1.
Fu-Anderton-Yang-Brauckmann does not explicitly disclose but Khalil discloses wherein the start of the portion of the multimedia content that is to be tagged is at a gap of a predefined time interval from the first time instance (Khalil, para. 0071, the user may provide input to enable or disable the variable length, and may also provide input indicating the number of frames or length of time that may be defined as variable start or end frames at the start and/or end of the customized video segment).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Khalil’s features into Fu-Anderton-Yang-Brauckmann’s invention for enhancing user’s video searching experience by supporting variable start and end point of segments to be tagged.
Regarding claim 5, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1, wherein the user can be an organizer of a live user interview (see Anderton-Yang, para. 0078).
Fu-Anderton-Yang-Brauckmann does not explicitly disclose but Khalil discloses receiving, by the application server, via an organizer device of an organizer of the live user interview, a second alert associated with the multimedia content at the second time instance, wherein the second alert is indicative of the end of the portion of the multimedia content that is to be tagged, and wherein the second time instance is determined by the application server based on the reception of the second alert (Khalil, para. 0071, receiving, from the user, input to enable or disable the variable length, and may also provide input indicating the number of frames or length of time that may be defined as variable start or end frames at the start and/or end of the customized video segment).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Khalil’s features into Fu-Anderton-Yang-Brauckmann’s invention for enhancing user’s video searching experience by allowing user to indicate end point of segments to be tagged.
Regarding claim 6, Fu-Anderton-Yang-Brauckmann-Khalil discloses the method of claim 5, wherein the second alert is received while the multimedia content is being recorded (Khalil, para. 0071, the user may provide input to enable or disable the variable length, and may also provide input during recording indicating the number of frames or length of time that may be defined as variable start or end frames at the start and/or end of the customized video segment).
The motivation to combine the references and obviousness arguments are the same as claim 5.
Regarding claim 7, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1.
Fu-Anderton-Yang-Brauckmann does not explicitly disclose but Khalil disclose wherein the second time instance is determined by the application server to be at a predefined time duration after the first time instance (Khalil, para. 0071, the user may provide input to enable or disable the variable length, and may also provide input indicating the number of frames or length of time that may be defined as variable start or end frames at the start and/or end of the customized video segment).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Khalil’s features into Fu-Anderton-Yang-Brauckmann’s invention for enhancing user’s video searching experience by providing predefined duration of segments to be tagged.
6. Claims 9-10 are rejected under AIA 35 U.S.C. 103 as being unpatentable over Fu-Anderton-Yang-Brauckmann, as applied to claim 1 above, in view of Schrantz et al. (US Publication 2022/0021717).
Regarding claim 9, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1, wherein the user can be an organizer of a live user interview (see Anderton-Yang, para. 0078), the method further comprises: storing, by the application server, the plurality of tags in the memory, wherein each tag of the plurality of tags is indicative of at least one context associated with the multimedia content; and receiving, by the application server, a context indicator that is indicative of the context of the portion of the multimedia content to be tagged, wherein the first tag is identified from the plurality of tags based on the context indicator (Fu, para. 0093, the above-mentioned content recognition of the recorded video to obtain n video segments may include: identifying, from the recorded video, the start and end times, i.e., “a first time instance and a second time instance”, of a song, applause for a joke, a message, a live chat, gift-giving, cheers, etc., extracting the video segment between the start and end times of the song, applause for a joke, a message, a live chat, gift-giving, cheers, etc., and obtaining video segment of a song, a joke, a message, a live chat, gift giving, victory, etc.; para’s 0097-0101, after obtaining the above n video segments, the tag information of each video segment can be obtained; the above tag information is used to characterize the features of video segments including content tags, which are used to indicate the content category “context indicator” corresponding to the video clip. For example, the content tag may include at least one of the following: song, joke, comment, gift, competition victory, talent; as such, the obtained/generated video segment includes video context, for example, the song, the joke, the message, the live chat, the gift giving, or the victory in the multimedia content that is to be tagged; Brauckmann, para’s 0007-0010, storing the tag in association with the video data, in particular in a header of a file comprising the video data or including a link between the tag and the video data. A tag is a piece of information that is related to the video data, in particular information related to an object of interest therein, such a keyword or a short text/description).
Fu-Anderton-Yang-Brauckmann does not explicitly disclose but Schrantz discloses receiving the context indicator via an organizer device of an organizer of the live user interview (Schrantz, para. 0086, FIG. 12 shows a content management interface (618) that may be used to view content session information and interact with content sessions (e.g., individual video interview sessions). Content session information shown may include a thumbnail of the content (620), a transcription indicator (622) that, when present, indicates that the content has been transcribed and transcription results are available, a duration and type (624) indicating the length of the content session and type of content (e.g., resolution of video content, bitrate of audio content), a number of clips (626) that have been created from a particular content session, an edit button (628) usable to navigate to an editing interface for a particular content session, a tag indicator (630) indicating a number of comments, tags, or keywords that a collaborating user associated with a content session (as described in the context of FIGS. 7 and 8), a filter input (631) that may receive text to be used to filter content sessions to those tagged with search terms or associated with transcription metadata containing search terms, and a transcription request button (632) usable to submit the content session for transcription. Each piece of information or button shown may be interacted with by a user to gain more information (e.g., hovering over a project thumbnail (620) to view a larger version), navigate to a new interface (e.g., clicking on the number of clips (626) to view a clip specific interface, clicking on the number of tags (630) to view a tag specific interface), or perform another action (e.g., clicking on the transcription button (632) to cause content to be transcribed).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Schrantz’s features into Fu-Anderton-Yang-Brauckmann’s invention for enhancing user’s video searching experience by generating video index using user input of context indicator that corresponds to the segment to be tagged.
Regarding claim 10, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1, further comprising: receiving, by the application server, a second tag for the first multimedia clip; linking, by the application server, the second tag with the first multimedia clip; and storing, by the application server, the first multimedia clip and the corresponding second tag in the memory associated with the application server (Fu, para. 0093, the above-mentioned content recognition of the recorded video to obtain n video segments may include: identifying, from the recorded video, the start and end times, i.e., “a first time instance and a second time instance”, of a song, applause for a joke, a message, a live chat, gift-giving, cheers, etc., extracting the video segment between the start and end times of the song, applause for a joke, a message, a live chat, gift-giving, cheers, etc., and obtaining video segment of a song, a joke, a message, a live chat, gift giving, victory, etc.; para’s 0097-0101, after obtaining the above n video segments, the tag information of each video segment can be obtained; the above tag information is used to characterize the features of video segments including content tags, which are used to indicate the content category corresponding to the video clip. For example, the content tag may include at least one of the following: song, joke, comment, gift, competition victory, talent; as such, the obtained/generated video segment includes video context, for example, the song, the joke, the message, the live chat, the gift giving, or the victory in the multimedia content that is to be tagged; as such, a second content tag associated with different context of the video clip may be obtained; Brauckmann, para’s 0007-0010, storing the tag in association with the video data, in particular in a header of a file comprising the video data or including a link between the tag and the video data. A tag is a piece of information that is related to the video data, in particular information related to an object of interest therein, such as a keyword or a short text/description; as such, a second content tag associated with different object(s) in the video clip may be obtained).
Fu-Anderton-Yang-Brauckmann does not explicitly disclose but Schrantz discloses receiving a second tag via an organizer device of an organizer of the live user interview for the first multimedia clip (Schrantz, para. 0086, FIG. 12 shows a content management interface (618) that may be used to view content session information and interact with content sessions (e.g., individual video interview sessions). Content session information shown may include a thumbnail of the content (620), a transcription indicator (622) that, when present, indicates that the content has been transcribed and transcription results are available, a duration and type (624) indicating the length of the content session and type of content (e.g., resolution of video content, bitrate of audio content), a number of clips (626) that have been created from a particular content session, an edit button (628) usable to navigate to an editing interface for a particular content session, a tag indicator (630) indicating a number of comments, tags, or keywords that a collaborating user associated with a content session (as described in the context of FIGS. 7 and 8), a filter input (631) that may receive text to be used to filter content sessions to those tagged with search terms or associated with transcription metadata containing search terms, and a transcription request button (632) usable to submit the content session for transcription. Each piece of information or button shown may be interacted with by a user to gain more information (e.g., hovering over a project thumbnail (620) to view a larger version), navigate to a new interface (e.g., clicking on the number of clips (626) to view a clip specific interface, clicking on the number of tags (630) to view a tag specific interface), or perform another action (e.g., clicking on the transcription button (632) to cause content to be transcribed).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Schrantz’s features into Fu-Anderton-Yang-Brauckmann’s invention for enhancing user’s video searching experience by providing video index using a plurality of tags that correspond to segments to be tagged.
7. Claims 12-13 are rejected under AIA 35 U.S.C. 103 as being unpatentable over Fu-Anderton-Yang-Brauckmann, as applied to claim 1 above, in view of Teller (US Patent 8,937,620).
Regarding claim 12, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1, wherein the user can be an organizer of a live user interview (see Anderton-Yang, para. 0078) and further disclose changing the labels that correspond to the recorded video (see Anderton-Yang, para’s 0062 and 0066, changing labels; Brauckmann, para’s 0007-0010, storing the tag in association with the video data, in particular in a header of a file comprising the video data or including a link between the tag and the video data. As known in the art, a tag and a label individually is a piece of information that is related to the video data, in particular information related to an object of interest therein, such as a keyword or a short text/description; as such, a label can be stored in conjunction with a video and a tag that corresponds to the video content; as such, a label can be stored in conjunction with a video and a tag that corresponds to the video content).
Fu-Anderton-Yang-Brauckmann does not explicitly disclose but Teller discloses receiving an input indicative of an instruction to delink the first tag from the first multimedia clip; and updating the first multimedia clip and the corresponding first tag stored in the memory to delink the first tag from the first multimedia clip (Teller, col. 14 lines 58-63, editing a tag may include changing a first tag associated with a content character to a second tag “delink the first tag”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Teller’s features into Fu-Anderton-Yang-Brauckmann’s invention for enhancing user’s video searching experience by effectively updating the video index based on updated tags related to relevant segments.
Regarding claim 13, Fu-Anderton-Yang-Brauckmann discloses the method of claim 1, wherein the user can be an organizer of a live user interview (see Anderton-Yang, para. 0078) and further disclose changing the labels that correspond to the recorded video (see Anderton-Yang, para’s 0062 and 0066, changing labels; Brauckmann, para’s 0007-0010, storing the tag in association with the video data, in particular in a header of a file comprising the video data or including a link between the tag and the video data. As known in the art, a tag and a label individually is a piece of information that is related to the video data, in particular information related to an object of interest therein, such as a keyword or a short text/description; as such, a label can be stored in conjunction with a video and a tag that corresponds to the video content; as such, a label can be stored in conjunction with a video and a tag that corresponds to the video content).
Fu-Anderton-Yang-Brauckmann does not explicitly disclose but Teller discloses receiving an input indicative of an instruction to delink the first tag from the first multimedia clip and link a second tag to the first multimedia clip; updating, by the application server, the first multimedia clip and the corresponding first tag stored in the memory to delink the first tag from the first multimedia clip; linking, by the application server, the second tag with the first multimedia clip; and storing, by the application server, the first multimedia clip and the corresponding second tag in the memory associated with the application server (Teller, col. 14 lines 58-63, editing a tag may include changing a first tag associated with a content character to a second tag “delink the first tag” and updating the first tag with the second tag. Storing the first multimedia clip and the corresponding second tag is known in the art, see Brauckmann as disclosed above, para’s 0007-0010, storing the tag in association with the video data, in particular in a header of a file comprising the video data or including a link between the tag and the video data. As known in the art, a tag and a label individually is a piece of information that is related to the video data, in particular information related to an object of interest therein, such as a keyword or a short text/description).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate Teller’s features into Fu-Anderton-Yang-Brauckmann’s invention for enhancing user’s video searching experience by effectively updating the video index based on updated tags that correspond to segments to be tagged.
8. The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. These include:
Zavesky et al., US Publication 2022/0167068
Saito et al., US Publication 2023/0362315
Sabo, US Patent 11,245,947
Conclusion
9. Any inquiry concerning this communication or earlier communications from the examiner should be directed to LOI H TRAN whose telephone number is (571)270-5645. The examiner can normally be reached 8:00AM-5:00PM PST FIRST FRIDAY OF BIWEEK OFF.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, THAI TRAN can be reached at 571-272-7382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/LOI H TRAN/Primary Examiner, Art Unit 2484