DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments filed August 24, 2026 have been fully considered but are not persuasive.
Regarding the rejection of claims 1, 3, 6-8 and 10-15 under 35 U.S.C. 103 as being unpatentable over Dareddy in view of Chen, Applicant argues that the combined teachings of Dareddy and Chen do not teach the combination of elements recited in amended independent claims 1, 6 and 7. These arguments are largely moot in view of the withdrawal of the rejection. However, a new ground of rejection is set forth below that does not rely on the teachings of Dareddy, but does rely on the teachings of Chen. Therefore, Applicant’s arguments regarding Chen are addressed below. Applicant’s arguments regarding Yerva are also addressed below since the new ground of rejection is based on the combined teachings of Yerva and Chen.
Regarding Chen, Applicant argues that Chen discloses two different training processes that are different from that of the present invention and argues that Chen does not teach the training features recited in claim 1. It should be noted that Chen was not relied on as teaching the training features of amended claim 1. Chen was relied on for its teaching of determining events corresponding to scene changes in a video such as a movie by extracting and processing auxiliary information comprising video shots that are a predetermined time period after the images corresponding to the scene change event to determine whether the current shot includes a scene change event.
Applicant argues that “Yerva describes detecting game events based on states or changes in UI information (e.g., cash or ammunition) across proximate frames. See Yerva ,i,i [0022], [0032], [0036]. But Yerva describes using such detected game events to select portions of highlight video - not for generating corrected data for model training.”
The examiner disagrees. Yerva describes using the detected events and the corresponding input region that caused the system to detect the event to train machine learning models to detect events. Para. [0036] discusses the manner in which the highlight generation system 500 shown in Fig. 5 uses machine learning to recognize events based on the input region images and the detected events, and also discloses that neural networks are trained to classify the events based on the input regions and the corresponding detected events. Para. [0037] also discusses training: “[a]s mentioned previously, identifying events can be useful for other purposes as well, such as for testing or training purposes.”
Para. [0038] of Yerva also discusses training: “[t]his can include, for example, attempting to determine one or more objects, actions, events, or occurrences that may be indicative of a game, or mode of gameplay. This data may be determined in at least one embodiment by using one or more neural networks, such as one or more convolutional neural networks (CNNs) trained to recognize different types of objects in image or video data.” Paras. [0042]-[0062] discuss, with reference to Figs. 7A and 7B, training neural networks to perform event recognition based on the extracted images and auxiliary information. The calculation of weight parameters that are adjusted during training is also discussed in this portion of Yerva. These portions of Yerva also discuss using the models to perform inferencing after they have been trained. Therefore, Yerva discloses generating corrected data, i.e, the detected event, for model training.
Claim Interpretation
The claims in this application are given their broadest reasonable interpretation (BRI) using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The BRI of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification.
In the following, some of the terms in the claims have been given BRIs in light of the specification. These BRIs are used for purposes of searching for prior art and examining the claims, but cannot be incorporated into the claims. Should Applicant believe that different interpretations are appropriate, Applicant should point to the portions of the specification that clearly support a different interpretation.
The term “input region” is interpreted as video content within the “target region” described with reference to Fig. 5 as the central region 70 of the displayed screen (para. [0033]).
The term “auxiliary information” is interpreted as corresponding to user interface (UI) elements of the displayed screen, such as first and second hit-point (HP) gauges 75 and 76, respectively, shown in Fig. 5, or as sound or audio data (paras. [0014], [0033], [0049]).
The phrase “correct data” is interpreted as information indicative of the occurrence of an event or information indicative of an event type (paras. [0034], [0050], [0051]).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 3 and 6-15 are rejected under 35 U.S.C. 103 as being unpatentable over U.S. Publ. Appl. No. 2022/0118363 A1 to Yerva et al. (hereinafter referred to as “Yerva”) in view of U.S. Patent No. 11,748,988 to Chen et al. (hereinafter referred to as “Chen”).
Regarding claim 1, Yerva discloses an image analysis system (highlight generation system 500, Fig. 5, para. [0036]) comprising:
extract an input region (Fig. 1, para. [0022], the extracted input region is the portion of image 100 that excludes the UI icons for player eliminated 102, time remaining 104, in-game chat messages 108, type of ammunition or weapon selected 110, amount of ammo remaining, shield 114, health 116, virtual player cash 118, and location 120) from video content including multiple time-series images and time-series audio data (Para. [0019] discusses the video content, which includes including multiple time-series images and time-series audio data: “[i]n at least one embodiment, this media content can include audio and video representative of one of these other types of experiences, such as streaming video of a gaming session of another player”), the input region being extracted from a first portion of each image of a group of the multiple time-series images of the video content (the input region is a first portion of image 100 that excludes the UI icons for player eliminated 102, time remaining 104, in-game chat messages 108, type of ammunition or weapon selected 110, amount of ammo remaining, shield 114, health 116, virtual player cash 118, and location 120), the input region including an object (Fig. 1A, the input region is the highlight event region that includes one or more objects, e.g., the first portion of the image 100 that includes the input region includes one or more objects, such as the spider and the player’s hand holding the weapon shown in Fig. 1A);
extract auxiliary information from a second portion of the video content (the auxiliary information includes the states of one or more of the UI icons for player eliminated 102, time remaining 104, in-game chat messages 108, type of ammunition or weapon selected 110, amount of ammo remaining, shield 114, health 116, virtual player cash 118, and location 120. The second portion is a peripheral region of the image 100 in which the UI icons are located), the second portion being at a predetermined time period after the first portion from which the input region is extracted (Para. [0022], as discussed in more detail below, a change in the status or state of the UI icons is used as auxiliary information that is processed to determine whether a highlight event has occurred. To determine whether there is a change in the state of one of the UI icons, the state corresponding to the icon extracted from a frame that occurred earlier in time is compared with the state of the icon extracted from a frame that occurred later in time. The change in state is used to determine whether a highlight event occurred: “[t]he information contained in at least some of these regions can change over time, and those changes can be indicative of various types of events. In at least one embodiment, events can be determined by detecting changes in one or more of these regions.” The corresponding event must occur before the status of the icon changes since the occurrence of the event causes the status change, and therefore the frame that contains the second portion corresponding to the changed status of the icon must occur a time period after the frame that contains the first portion that contains the corresponding highlight event. For example, if one player is killed by another player, the frame that contains the changed state of the player-eliminated icon 102 will occur a time period after the frame that contained the event of the player being killed. Yerva does not explicitly disclose that this known time period is a predetermined time period);
processing the extracted auxiliary information to:
obtain, from the extracted auxiliary information, a change in a value of a parameter associated with the object based on an image that precedes, by a predetermined number of frames, the image from which the auxiliary information is extracted (as indicated above, a change in the value of the state of the icon, which is a “parameter associated with the object”, is obtained from the auxiliary information over multiple frames to obtain the change that is used to determine whether a highlight event has occurred. The change is obtained based on an image contained in the frame at the time of, or shortly before, the event and based on the image contained in the frame that occurred later in time that includes the extracted auxiliary information that indicates the changed state. The number of frames of separation between these two frames depends on the configuration of the system, but at least two frames are needed. While Yerva does disclose analyzing frames of video content to detect changes between frames (paras. [0021], [0023] and [0028]), Yerva does not explicitly disclose that there is a predetermined number of frames between the frame that occurred before the state value changed and the frame that contains the state value change);
detect, based on the change in the value of the parameter, an event for the object included in the input region in the auxiliary information and generating correct data as the detected event (As indicated above, the occurrence of an event is detected based at least in part on the change of the state value of the icon. In Yerva, the “correct data” corresponds to the recognized event. Para. [0022]: “[t]here may be various other regions that correspond to graphical user interface (GUI) or heads-up display (HUD) information as well, which may be useful in identifying these and other types of events that occur during gameplay. For example, image 100 includes regions that correspond to various UI elements, as relate to time remaining 104, in-game chat messages 108, type of ammunition or weapon selected 110, amount of ammo remaining, shield 114, health 116, virtual player cash 118, and location 120. There may be other regions associated with information that only appears at certain times, such as when a player dies and is spectating gameplay of another player. The information contained in at least some of these regions can change over time, and those changes can be indicative of various types of events. In at least one embodiment, events can be determined by detecting changes in one or more of these regions, and combining information for that change with information in one or more other reasons [sic] that may be used to determine a type of event that has occurred.”);
generate correct data indicating the detected event for the object (in Yerva, correct data is generated corresponding to the type and timing of the event, para. [0036]: “[a]n event recognition module can use one or more event auto-recognition algorithms, processes, or deep learning approaches to recognize events, or objects and occurrences associated with various types of events…the event recognition module 506 can evaluate one or more subordinate regions, as may be determined by one or more rules for one or more specific types of event. Information from these subordinate regions can then be passed to the event analysis module 508 to determine whether one or more highlight selection criteria have been satisfied. For a kill event, a highlight selection criterion might include a determination that a current player killed another player with at least 85% certainty based at least in part upon the information from these regions. If such a criterion is satisfied, information for that event can be passed to a highlight generation module 510, which can be responsible for generating a corresponding highlight. This can include, for example, pulling video data from a video buffer 514, where the video may include some amount of video content before, and after, a timing of the event. In another embodiment, this may include determining timing information for this highlight to be used to pull that video content at a later time. This highlight information for one or more highlights 512 can then be provided as output, to be stored for subsequent viewing or presentation via a client device as those highlights are determined.”) ; and
train a machine learning model based on training data including the input region and the correct data, wherein the input region is extracted from an image earlier than a timing at which the correct data is generated, wherein the machine learning model is trained to detect events in the video content based on the training data, wherein the training comprises (Yerva discloses using the input region of the image that contained the highlight event and the correct data corresponding to the type of event that was recognized to train a machine learning model. Para. [0036] discusses the manner in which the highlight generation system 500 shown in Fig. 5 uses machine learning to recognize events based on, for example, the input region containing the event and the event recognition result. This portion of Yerva also discloses that neural networks are trained to classify the events. Para. [0037] also discusses training: “[a]s mentioned previously, identifying events can be useful for other purposes as well, such as for testing or training purposes.” Para. [0038] of Yerva also discusses training: “[t]his can include, for example, attempting to determine one or more objects, actions, events, or occurrences that may be indicative of a game, or mode of gameplay. This data may be determined in at least one embodiment by using one or more neural networks, such as one or more convolutional neural networks (CNNs) trained to recognize different types of objects in image or video data.” Paras. [0042]-[0062] discuss, with reference to Figs. 7A and 7B, the training of neural networks to perform event recognition based on the extracted images and auxiliary information. The calculation of weight parameters of the neural network that are adjusted during training is also discussed. These portions of Yerva also discuss using the models to perform inferencing after training them):
processing the input region using the machine learning model to generate output information indicative of occurrence of an event for the object (Paras. [0042]-[0051] discuss inferencing/training logic 715 used to perform inferencing by a neural network to identify highlight events and to perform training of the neural network to perform the inferencing. Paras. [0079]-[0080] discuss inferencing by the inference and training logic 715 to determine an occurrence of a highlight event); and
adjusting one or more learning parameters of the machine learning model based on (i) the output information generated from the input region extracted from the first portion and (ii) the correct data generated, based on the change in the value of the parameter associated with the object, from the auxiliary information extracted from the second portion (The BRI for this limitation is that parameters of the machine learning model are adjusted, e.g., via back propagation, based on the event occurrence identified by the machine learning model and based on change of value of the parameter that led to the machine learning model identifying the event. Paras. [0036]-[0038] disclose using the identified event type and the data that led to identifying the event type, such as the detected change in state, to train a neural network to classify such events. Para. [0043] discusses weights and other parameters of the layers of the neural network. Para. [0045] discusses adjusting weights via backward propagation).
As indicated above, Yerva discloses that the second portion of the video content from which the auxiliary information is extracted is some time period after the first portion from which the input region is extracted, but does not disclose that the time period is predetermined. Chen, in the same field of endeavor, discloses determining events corresponding to scene changes in a video such as a movie by extracting and processing auxiliary information comprising video shots that are a predetermined time period after the images corresponding to the scene change event to determine whether the current shot includes a scene change (Chen, Col. 2, line 1 – Col. 3, line 23, discusses processing shots of video sequences to determine whether a current shot contains a scene change, or boundary; Fig. 4A, Col. 9, line 65 – Col. 10, line 28 discloses using features of shots 6 and 7 as auxiliary information to determine whether the current shot 5 includes a scene change, where shots 6 and 7 each comprise a sequence of image frames that are a predetermined time period after the sequence of image frames comprising shot 5, where the predetermined time period is equal to the distance in time between shot 5 and shots 6 and 7 and on the time duration of the shots).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present disclosure to modify the system of Yerva to extract auxiliary information that is located a predetermined time period after the images from which the input region video is extracted as taught by Chen. A person of ordinary skill in the art would have been motivated to make the modification to improve the accuracy of event detection by using not only video frames corresponding to the current frame to detect events, but also video frames that follow the current frame in time by a predetermined time period. The modification could have been made by a person of ordinary skill in the art with a reasonable expectation success because making the modification merely involves combining prior art elements according to known methods to yield predictable results (modifying software executed by the highlight generation system 504 shown in Fig. 5 to implement the predetermined timing control of when the second portions containing the auxiliary information are extracted for use in determining state value changes).
As indicated above, although Yerva does disclose analyzing frames of video content to detect changes between frames (paras. [0021], [0023] and [0028]), Yerva does not explicitly disclose that there are a predetermined number of frames between the frame that included the image of the icon before the state value changed and the frame that included the image of the icon after the state value changed. In Yerva, the change in state value is obtained based on an image of the icon contained in a frame that occurred at the time of, or before, the frame that contained the event, and based on the image of the icon contained in a frame that occurred sometime after the event that resulted in the state change. Yerva does not explicitly disclose that the frames that are used to determine the change of state are separated by a predetermined number of frames.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present disclosure to try different predetermined numbers of frames to determine which number results in the most accurate results since there are a finite number of choices that will work with predictability. In order for the highlight event to be captured in time, the system of Yerva identifies the frame in which the event occurred by using frames that are separated in time from one another, such as a frame in which there is no state value change and the next frame in time in which there is a state value change. If the predetermined number is too large, then the highlight event that corresponds to the change may not be accurately captured and more memory may be needed to store the highlight event clip. Therefore, the predetermined number that is used for this purpose is a design choice that can be selected with predictability based on considerations such as the amount of memory that is needed to store the highlight clips and the desired lengths of the highlight clips. (See KSR International Co. v. Teleflex Inc., 550 U.S. 398, 82 USPQ2d 1385 (2007), in which the Court held that "obvious to try" was a valid rationale for an obviousness finding, for example, when there is a "design need" or "market demand" and there are a "finite number" of solutions. 550 U.S. at 421, 82 USPQ2d at 1397).
Regarding claim 3, as indicated above, Yerva does not explicitly disclose that the predetermined time period is defined as associated with a predefined number of consecutive images from the video content. As indicated above in the rejection of claim 1, Chen discloses extracting the auxiliary information (video shots that follow the current video shot in time) from a portion of the video content that is located at a predetermined time period following the current shot. Chen defines a shot as a contiguous sequence of video frames (Col. 2, lines 16-23). In Chen, the predetermined time period is based on the number of frames back that the shots 6 and 7 are in time from the current shot 5. In the example described in Chen with reference to Fig. 4A, features of shots 6 and 7 are used to determine whether current shot 5 contains a scene change. In that example, the predetermined time period is based on the length of time over which each shot extends (i.e., the length of the contiguous sequence of frames comprising the shot) and how far back in time the shots are that are used as auxiliary information. Therefore, the predetermined time period in Chen corresponds to a predetermined number of consecutive images that the following shots are from the current shot, which would be based on predetermined system settings.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present disclosure to modify the listening device 500 of Dareddy based on the teachings of Chen to extract auxiliary information comprising audio that is located a predetermined number of consecutive image frames after the current shot comprising the video input region as taught by Chen. A person of ordinary skill in the art would have been motivated to make the modification to improve the accuracy of event detection by using not only audio frames corresponding to the current frame, but also audio frames that follow the current shot in time by a consecutive predetermined number of frames. The modification could have been made by a person of ordinary skill in the art with a reasonable expectation success because making the modification merely involves combining prior art elements according to known methods to yield predictable results (modifying the software comprising learning module 504b to ensure that the audio frames that are processed to detect the occurrence of a highlight event include audio frames that follow the group of images from which the input region is extracted by a predetermined amount of time).
Regarding claim 6, the rejection of claim 1 above applies mutatis mutandis to claim 6.
Regarding claim 7, to the extent that claim 7 recites the same limitations that are recited in claim 1, the rejection of claim 1 above applies mutatis mutandis to claim 7. The only limitations that are recited in claim 7 that are not also recited in claim 1 is the non-transitory computer-readable medium storing instructions for performing the method recited in claims 1 and 7. Yerva discloses a non-transitory computer-readable medium storing instructions for performing the method (Para. [0151]).
Regarding claim 8, Yerva discloses that the second portion of the video content from which the auxiliary information is extracted is an audio portion (Para. [0021] discloses that analyzing the video content to detect highlight events can include analyzing an audio portion. Yerva discloses that “at least one audio region may be specified as well, as may relate to a sound or music that plays in response to, or along with, a type of event.”).
Regarding claim 9, as indicated above, Yerva, a portion of which is duplicated below for convenience, discloses that the correct answer, i.e., detecting that a particular event has occurred, can be generated on the basis of a change in the UI elements 102-120 shown in Fig. 1A. (Para. [0022] of Yerva: “[t]he information contained in at least some of these regions can change over time, and those changes can be indicative of various types of events. In at least one embodiment, events can be determined by detecting changes in one or more of these regions, and combining information for that change with information in one or more other reasons that may be used to determine a type of event that has occurred”; see also Figs. 5A and 5B and para. [0036]: “[i]n this example, the event recognition module 506 can analyze information in these primary regions, and can pass this information to an event analysis module 508. An event recognition module can use one or more event auto-recognition algorithms, processes, or deep learning approaches to recognize events, or objects and occurrences associated with various types of events.”).
Because the UI elements 102-120 that are used in Yerva as the auxiliary information are disposed around the periphery of the main scene, but are part of the video content, the second portion of the video content from which the UI element images are extracted is a different region with images of the video content than the video content of the input region, which would correspond to the primary image of the scene shown below in Fig. 1A of Yerva that does not include the UI elements.
PNG
media_image1.png
200
400
media_image1.png
Greyscale
Regarding claim 10, the rejection of claim 3 above applies mutatis mutandis to claim 10.
Regarding claim 11, the rejection of claim 8 above applies mutatis mutandis to claim 11.
Regarding claim 12, the rejection of claim 9 above applies mutatis mutandis to claim 12.
Regarding claim 13, the rejection of claim 3 above applies mutatis mutandis to claim 13.
Regarding claim 14, the rejection of claim 8 above applies mutatis mutandis to claim 14.
Regarding claim 15, the rejection of claim 9 above applies mutatis mutandis to claim 15.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DANIEL J SANTOS whose telephone number is (571)272-2867. The examiner can normally be reached M-F 9-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matt Bella can be reached at (571)272-7778. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DANIEL J. SANTOS/Examiner, Art Unit 2667
/MATTHEW C BELLA/Supervisory Patent Examiner, Art Unit 2667