DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 5-9, 12-16, and 19-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lin (KR 20210114074 A) in view of Jindal (PGPUB: 20210103615 A1), and further in view of (PGPUB: 20200084519 A1).
Regarding claims 1, 9, and 16. Lin teaches a method for generating a video caption, comprising:
obtaining a video to be processed (see page 7, lines 29-32, the frames shown in FIG. 3 are selected from video (ellipses in the figures are omitted to indicate frames not shown), and each frame is individually processed by a CNN encoder to obtain a picture of each selected video frame);
identifying a plurality of keyframes corresponding to the video to be processed (see page 10, lines 3-6, for a given video, for example, selecting frames at the same interval from the given video, or using a keyframe algorithm in which frames can be selected via a neural network, multiple image frames (i.e., keyframes) from a given video. ) can be selected), and auxiliary caption information (see page 10, lines 14-16, the local visual feature fully uses the inter-frame information of each image, which can create more accurate text captions of videos or images; see page 29, lines 1-6, a text caption of a video obtained on the basis of graph convolution characteristics (which as can be understood may also be spatio-temporal visual characteristics and/or semantic characteristics and graph convolution characteristics) For , a vector relating to the graph convolution feature may be obtained. These feature vectors are input to a self-attention-based decoder to learn to obtain inter-frame information between frames),
wherein the keyframes include a main object, and the auxiliary caption information includes at least one of the following (see page 11-12, lines 37-40 and 1-3, a scene graph refers to a graph structure in which local visual characteristics, attribute characteristics (to be described in detail later), and relational characteristics of respective target regions in the image are expressed in a graph manner. The scene graph may include multiple nodes and multiple edges, where each node of the multiple nodes is a target feature (ie, a local visual feature above) or a property of a target (ie an object) included in the target area):
name information corresponding to the main object, object category corresponding to the main object (see page 23, lines 24-29, for different types of multimedia data (such as different videos and different images), the importance of the characteristic information of each classification is likely to be different, and since the characteristic information of different classifications may have different weights, Different features play different roles, and thus, the solution of an embodiment of the present disclosure may be adaptable to generating captioning information of different videos), object attribute corresponding to the main object (see page 15, lines 29-32, different classifiers can be used to obtain different types of attributes. In this way, the acquired properties are more accurate and the corresponding properties are more diversified to generate more accurate captioning information based on the predicted property characteristics); and
auxiliary features corresponding to the auxiliary caption information (see page 11, lines 29-36, based on the extracted characteristic information, generating the text caption of the multimedia data includes: acquiring relational characteristics between targets based on local visual characteristics of each target in the image; building a scene graph of the image based on the local visual features and the relationship features between the targets; obtaining graph convolution features of the image based on a scene graph of the image; and generating a text caption of the multimedia data based on graph convolution characteristics of each image of the multimedia data);
generating, based on the image features and the auxiliary features, a target caption corresponding to the video to be processed (see page , lines , Extracting local visual features of targets included in respective target regions of each image in the multimedia data to do; extracting semantic features of the multimedia data; when the multimedia data is a video, extracting spatial-temporal visual features of the multimedia data; extracting global visual features of the multimedia data; extracting attribute characteristics of the targets included in respective target regions of each image in the multimedia data; and extracting global attribute features of each image from the multimedia data. That is, the characteristic information of the multimedia data may include at least one of a local visual characteristic, a semantic characteristic, a spatial-temporal visual characteristic, a global visual characteristic, a local attribute characteristic (ie, an attribute characteristic of a target), and a global attribute characteristic. have. The local visual feature is a visual feature of a target area within an image, ie, the target area is relatively local to the image to which the target area belongs),
Lin does not expressly teach:
video tag corresponding to the video to be processed, and
determining an image feature corresponding to each of the plurality of keyframes, and auxiliary features corresponding to the auxiliary caption information;
generating, based on the image features and the auxiliary features, a target caption corresponding to the video to be processed,
wherein the target caption includes the name information of the main object.
Jindal teaches:
video tag corresponding to the video to be processed (see Fig. 1, paragraph 25, the image processing system 114 is used to detect keyframes, generate content tags for keyframes, and use the content tags to service search queries for video content), and
determining an image feature corresponding to each of the plurality of keyframes (see Fig. 1, paragraph 25, the image processing system 114 is used to obtain, detect, and identify features associated with one or more frames in a video file, where each frame is an image from the sequence of images in the video file. The image processing system 114 can include one or more processing devices for executing suitable program code for performing one or more functions. Examples of this program code include the software engines depicted in FIG. 1, such as keyframe detector 116, tag generator 118, and aesthetics engine 120. Image processing system 114 can use one or more of these engines to determine content tags for video content, select keyframes of the video content having content tags that match at least one keyword from a search query, and provide the selected keyframes to a rendering engine 122)
wherein the target caption includes the name information of the main object (see Fig. 4, paragraph 95, similar to block 204 of process 200, the image processing system 114 executes program code described herein (e.g., adaptive search engine 108, keyword search engine 110, keyframe detector 116, tag generator 118, etc.) to identify metadata associated with an image within a video frame (e.g., a creator, creation location, creation time, brand name, image title, one or more captions, keywords, technical details, digital rights, or any combination of these). In this example, the keyframe 400 depicts various features including jungle 402, sunshine 404, elephant 406, and water 408. In some embodiments, the tag generator 118 creates content tags 410 from a subset of the above-mentioned keyframe features).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination by Jindal to obtain to detect keyframes, generate content tags for keyframes, and use the content tags to service search queries for video content, in order to provide video tag corresponding to the video to be processed; further to obtain to obtain, detect, and identify features associated with one or more frames in a video file, where each frame is an image from the sequence of images in the video file and keyframe detector 116, tag generator 118, and aesthetics engine 120, in order to further provide determining an image feature corresponding to each of the plurality of keyframes; further to obtain process 200, the image processing system 114 executes program code described herein (e.g., adaptive search engine 108, keyword search engine 110, keyframe detector 116, tag generator 118, etc.) to identify metadata associated with an image within a video frame (e.g., a creator, creation location, creation time, brand name, image title, one or more captions, keywords, technical details, digital rights, or any combination of these), as wherein the target caption includes the name information of the main object. Therefore, combining the elements from prior arts according to known methods and technique would yield predictable results.
However, the combination does not expressly teach:
voice information corresponding to the video to be processed.
Pappu teaches to provide automatic labeling of video content with one or more tags that are directed to the semantic understanding of the video content as opposed or in addition to the simplistic identification of the objects, images, and/or sounds presented as part of the video content. Accordingly, the system and/or methods may label video content with tags that relate to the topics, content, and/or subject matter underlying the objects, images and/or sounds presented as part of the video content (see paragraph 11); the multimodal multilabel tagging may include partitioning (at 1) video content 110 into video modality 120, text modality 130, audio modality 140, and/or other modalities. Video modality 120 may include the visual or graphical elements of video content 110. Text modality 130 may include the speech, dialog, and/or text of video content 110. Text modality 130 may be obtained from closed captions that are associated with video content 110 as metadata or other data. Text modality 130 may alternatively be obtained via automated speech recognition and transcription. Audio modality 130 may include the audio of video content 110. More specifically, audio modality 130 may include the non-speech sounds and/or sound characteristics of video content 110 (see Fig. 1, paragraph 16).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the combination by Pappu to obtain Text modality 130 may include the speech, dialog, and/or text of video content 110. Text modality 130 may be obtained from closed captions that are associated with video content 110 as metadata or other data. Text modality 130 may alternatively be obtained via automated speech recognition and transcription. Audio modality 130 may include the audio of video content 110, in order to provide voice information corresponding to the video to be processed. Therefore, combining the elements from prior arts according to known methods and technique would yield predictable results.
Regarding claims 5, 12, and 19. The combination teaches the method according to claim 1, wherein determining the image feature corresponding to the image to be processed comprises:
segmenting the image to be processed to obtain a plurality of image blocks (see Fig.1, paragraph 44, the keyframe detector 116 creates video segments (e.g., temporal segments, topical segments, clustered segments, multi-scale segments, motion segments, chunks, etc.) from the features, where each video segment includes frames with a common set of features and thereby depicts a major scene from the video file);
determining a positional encoding corresponding to each of the plurality of image blocks (see Pappu, Fig. 1, paragraph 15, Video content 110 may include a set of video frames, audio, and/or metadata that are encoded as a file (e.g., an mp4 file), or a set of files (e.g., transport stream segments) with each file of the set of files encoding a different segment of video content 110. Video content 110 may span a duration of seconds, minutes, or hours. Each segment of video content 110 may span a subset of the overall duration);
processing the plurality of image blocks based on their respective positional encodings to obtain the image feature (see Pappu, Fig. 1, paragraph 31, a first segment of the news clip (e.g., first duration of two minutes) may cover a congressional race in a one state, a second segment of the news clip (e.g., second duration spanning a subsequent three minutes) may cover new legislation passed by Congress, and a third segment of the news clip (e.g., third duration spanning a final minute of the news clip) may cover a presidential press briefing. Each segment may be locally labeled with different tags to better identify the relevance of each segment, and also to allow for searching or indexing within video content 110).
Regarding claims 6, 13, and 19. The combination teaches the method according to claim 1, wherein when the auxiliary caption information does not include the object category corresponding to the main object, the method further comprises:
obtaining the object category of the main object in the image to be processed based on the image feature and the auxiliary feature (see Jindal, paragraph 18, a video file included in the search results could include keyframes that are related to the search query. The image processing system can compute, for each keyframe, a respective matching score. A matching score for a keyframe indicates a number of matches between content tags for the keyframe and a keyword set in the search query. Examples of search terms that could be included in a keyword set include one or more of a keyword specified by a user, a synonym of a user-specified keyword, a root word of a user-specified keyword (e.g., the root word “ride” for the user-specified word “riding”), an umbrella term encompassing a user-specified keyword (e.g., umbrella terms such as “primate” or “mammal” for the user-specified term “monkey”), a category associated with a user-specified keyword (e.g., categories including “funny monkey videos,” “cute baby monkeys,” or “monkey dancing video clips” for the user-specified term “monkey”), a semantically-related word or phrase for a user-specified keyword (e.g., semantically words or phrases such as “primate,” “ape,” “chimpanzee,” “chimp,” “species of great apes,” “new world monkeys,” or “old world monkeys” for the user-specified term “monkey”));
performing image classification based on the object category (see Lin, page 15, lines 13-15, the attribute prediction network may include multiple attribute classifiers, where each classifier corresponds to one type of attribute prediction) and the name information of the main object (see Fig. 4, paragraph 59, tag generator 450 may determine topical tags for video content 110 by matching extracted features from different modalities of video content 110 to tags from tag taxonomy 420 that are associated with the same or related features of video content in dataset 410. For instance, extracted features related to hotels, cuisine, and activities may be associated with a travel tag, and extracted features providing scores and various team names may be associated with a first tag for sports and a second tag for a particular sport).
Regarding claims 7 and 14. The combination teaches a non-transitory computer-readable storage medium configured with instructions executable by one or more processors to cause the one or more processors to perform the method of claim 1 (see claim 9 above).
Regarding claims 8 and 15. The combination teaches an electronic device comprising:
one or more processors; and one or more computer-readable memories coupled to the one or more processors and having instructions stored thereon that are executable by the one or more processors to perform the method of claim 1 (see claim 9 above).
Allowable Subject Matter
Claims 2-4, 10-11, and 17-18 objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to XIN JIA whose telephone number is (571)270-5536. The examiner can normally be reached 9:00 am-7:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Gregory Morse can be reached at (571)272-3838. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/XIN JIA/Primary Examiner, Art Unit 2663