DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 4/16/2026 has been entered.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 4-11, and 14-24 is/are rejected under 35 U.S.C. 103 as being unpatentable over Du et al1 (“Du”) in view of Abbas2 (“Abbas”).
Regarding claim 1, Du teaches a system comprising (note the system is addressed below with it comprising the apparatus configured to perform the operations in the body of the claim as addressed below): at least one processor; at least one memory component storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising (see Du, paragraph 0007 teaching “the techniques described herein relate to a computing device, including: a computing device display; a processor; and a memory configured with instructions to cause the processor to: generate an in-browser camera list, the in-browser camera list being selectable by a user within a browser; and add a virtual camera to the inbrowser camera list, the virtual camera being operable to receive a physical camera frame from a physical camera listed in the in-browser camera list and generate a modified frame based on a local browser setting” such that here the same components are utilized to perform the functions as will be further addressed below):
detecting selection of a virtual camera added as an extension in an internet browser (see Du, paragraphs 0018-0021 teaching “generating an in-browser camera list that is selectable by a user. The in-browser camera list may include a list of cameras, for example a list of physical cameras internal or external to a computing device, received from an operating system kernel to which the virtual camera has been added” where “the virtual camera may be installed for use within a browser via a plugin, add-on, or extension. Plugins, add-ons, and extensions are software additions that allow for the customization of web browsers” such that here the system is able to detect selection of a virtual camera which is added as an extension in a web browser);
generating a new canvas element for the virtual camera (note that a “generating a canvas element” is extremely as the manner of generation is not limited, nor specifically defined, and further neither “canvas” nor “element” is specifically limited and when interpreted as a functional element such a canvas is considered to be any drawable area or render target and is any functional element that can be drawn to and can be any type of surface to which something can be addressed – then it is necessary to interpret “canvas element” considered as recited together as a canvas element could be the canvas itself or could also be any element relating to a canvas such that some setting or function that defines, changes or modifies a canvas could be considered a canvas element; thus see Du, paragraph 0018-0021 teaching “Upon receiving an indication that a browser is accessing a virtual camera from the in-browser camera list, the method may include receiving a physical camera frame from a physical camera associated with the virtual camera. Next a modified frame may be generated based on the physical camera frame and a local browser setting” and “Next the modified frame may be sent displayed in the browser display and/or sent to another user participating in the video conference via another client application” and “result is that the modified frames from the virtual camera are seamlessly substituted for the physical camera feed” where here selecting a virtual camera generates a new canvas element for the system, as previously the cameras in the in-browser camera list provided the frames or surfaces or canvases for display, and thus selection of the virtual camera instead causes the canvas element to be the new surface corresponding to the virtual camera frames which take in a physical video frame feed and use this as a new canvas element for the virtual camera allowing the addition of selected AR effects or other filter effects, and thus the output area of the virtual camera is the new canvas element that is created and for example note that “the modified frame may be sent displayed in the browser display and/or sent to another user participating in the video conference via another client application” and “virtual camera receives frames from a physical camera, modifies those physical camera frames, and outputs the modified frames. The modified frames may be displayed, saved, or sent to another computing device” means that the canvas element being written to corresponds to the modified frame which is not actually displayed until sent to a display and for example the modified image frames could simply be “saved” such that again this means the modified images correspond to the new canvas element being drawn to where such canvas element may then be displayed – see further paragraphs 0022-0024 and figure 1A where “virtual camera display” corresponds to the new canvas element created for the virtual camera such that the result of displaying the modified image frame after being sent for display is “virtual camera display 106” which also may be considered to be a new canvas element that is generated each time the image is rendered as this is a canvas for display as well);
accessing video frames captured by a hardware camera coupled with a computing device (see Du, paragraphs 0018-0020 teaching “method may include receiving a physical camera frame from a physical camera associated with the virtual camera. Next a modified frame may be generated based on the physical camera frame and a local browser setting” such that here the system accesses the physical hardware camera coupled with the computing device);
detecting a selection of an augmented reality (AR) option (see Du, paragraphs 0018-0020 teaching “method may include receiving a physical camera frame from a physical camera associated with the virtual camera. Next a modified frame may be generated based on the physical camera frame and a local browser setting” and “local browser setting may be selectable by a user and enable features such as filters, animations, visual text interpretations, and so forth” where this local browser setting that augments the reality of the user is selected and is detecting of a selection of an AR option and this is what modifies the physical camera feed where paragraphs 0045-0048 provide additional options for selection of AR settings to apply such as “filter effects may comprise changing the pixels of a physical camera frame in an algorithmic way. For example, filter effects may comprise changing colors, creating brush or sketch effects, blurring, sharpening, brightening, changing the contrast, overlaying effects (e.g., a sunburst effect), etc. Sketch filter settings 112 in virtual camera setup 102 is one example of local browser setting 258” and “local browser setting 258 may relate to one or more animation effects. For example, animation effects may include generating an avatar, animal, or fanciful character (e.g., a robot character) that mimics the facial expression of the user. An example local browser setting 258 relating to an animation may comprise a drop-down box selecting an animal or character” and “text effects may include displaying a text interpretation of speech in the modified frame, similar to text effect 116 from FIG. IB, described above. The text interpretation may be generated from speech recorded in an audio track associated with the physical camera frames. In examples, text effects may include translating or summarizing speech. In examples, text effects may include animating the text to express context in the speech. In examples, the text effects may include moving the text interpretation around virtual camera display virtual camera display 107 in response to the first user moving. An example local browser setting 258 for a text effect may include selecting a language for translation, for example” such that these are all examples of AR options that can be selected from and detected in order to produce the modified image frame on the canvas element);
determining a number of video frames per second for applying the AR option to the accessed video frames based on a maximum rate of processing capability of the computing device and a length of time it takes to render the AR option;
applying the AR option to each video frame of a subset of the accessed video frames (see Du, paragraphs 0018-0020 teaching “a modified frame may be generated based on the physical camera frame and a local browser setting. The local browser setting may be selectable by a user and enable features such as filters, animations, visual text interpretations, and so forth. Next the modified frame may be sent displayed in the browser display and/or sent to another user participating in the video conference via another client application. The result is that the modified frames from the virtual camera are seamlessly substituted for the physical camera feed” and as in paragraphs 0045-0048 as explained above the selected AR option is applied which modifies the physical image of the camera feed in the virtual camera canvas element to create a modified image according to whatever the AR option selected was) at a predefined rate, based on the determined number of video frames per second for applying the AR option (see Du, paragraphs 0018-0020 as explained above where the AR option selected is applied at a predefined rate that matches the physical camera feed as “the modified frames from the virtual camera are seamlessly substituted for the physical camera feed” such that it is applied at a rate predefined to substitute each frame of the physical camera feed where “virtual camera receives frames from a physical camera, modifies those physical frames, and outputs the modified frames” where this means that from the perspective of applying the AR effect by the virtual camera, the predefined rate of applying is by selecting the subset of the accessed video frames at a 1:1 rate based on a predefined number of video frames per second where this predefined number of video frames per second corresponds to whatever rate that the physical camera is supplying as the physical camera feed is providing the frames at some rate which is predefined from the perspective of the virtual camera and then each of these frames is selected as the subset to which the AR effect is applied; see further paragraph 0026 teaching, “Once a browser is configured to use a virtual camera, the first user may join a video conference substituting the virtual camera feed for the physical camera feed” such that as the modified image effects are applied to each frame from the physical camera then they are applied at whatever the predefined rate of capture and display the video camera necessarily has such as for example any of the physical cameras and at this point each of the video frames is a subset of the accessed video frames as the frames have been accessed prior to joining of the video conference;);
providing each video frame of the subset of the accessed video frames comprising the applied AR option to the new canvas element (see Du, paragraphs 0018-0020 teaching “a modified frame may be generated based on the physical camera frame and a local browser setting. The local browser setting may be selectable by a user and enable features such as filters, animations, visual text interpretations, and so forth. Next the modified frame may be sent displayed in the browser display and/or sent to another user participating in the video conference via another client application. The result is that the modified frames from the virtual camera are seamlessly substituted for the physical camera feed” and as in paragraphs 0045-0048 as explained above the selected AR option is applied which modifies the physical image of the camera feed in the virtual camera canvas element to create a modified image according to whatever the AR option selected was and the so modified video image frame is provided as the ”modified image” and this element that is “sent [to be] displayed” is thus a provided video frame with the AR element option applied provided to the canvas element which is sent to be displayed, or again as noted above the display target that renders the modified image may also be considered provided such video frames to the new canvas element where this display target is the new canvas element); and
causing display, of each video frame of the subset of the accessed video frames comprising the applied AR option in the new canvas element, on a user interface of the computing device (see Du, paragraphs 0018-0020 teaching “a modified frame may be generated based on the physical camera frame and a local browser setting. The local browser setting may be selectable by a user and enable features such as filters, animations, visual text interpretations, and so forth. Next the modified frame may be sent displayed in the browser display and/or sent to another user participating in the video conference via another client application. The result is that the modified frames from the virtual camera are seamlessly substituted for the physical camera feed” and as in paragraphs 0045-0048 as explained above the selected AR option is applied which modifies the physical image of the camera feed in the virtual camera canvas element to create a modified image according to whatever the AR option selected was and the so modified video image frame is provided as the ”modified image” and this element that is “sent [to be] displayed” is thus a provided video frame with the AR element option applied provided to the canvas element which is sent to be displayed, or again as noted above the display target that renders the modified image may also be considered provided such video frames to the new canvas element where this display target is the new canvas element).
Du teaches all of the above, but fails to teach or suggest determining a number of video frames per second for applying the AR option to the accessed video frames based on a maximum rate of processing capability of the computing device and a length of time it takes to render the AR option and that in applying the AR option to each video frame of a subset of the accessed video frames that it is at a predefined rate based on the determined number of video frames per second for applying the AR option. Rather Du teaches to determine a number of video frames per second for applying the AR option by determining to apply the AR option to each video frame received which it then processes and uses to modify the AR experience based on the video frames and does not teach applying the AR option to a subset of the accessed video frames at a predefined rate based on that determined number of video frames per second for applying the AR option. Thus Du stands as a base device upon which the claimed invention can be seen as an improvement through such determining and applying steps which would allow the system to apply the AR processing options to only a subset of accessed video frames in response to rendering conditions and resources that may change such that the processing would more efficiently be performed with respect to the actual rendering environment, conditions, and resources.
In the same field of endeavor relating to rendering content in an augmented reality environment in which video frames captured by a hardware camera are processed to provide augmented reality content (see Abbas, column 10, lines 57-65 through lines 1-25 teaching “user video interfacing module 135 detects an interaction or user gesture event in connection with an act by the user during the time period based on at least one video frame of the live multimedia stream received during the time period. In some implementations, the user gesture event may correspond to an indication of a set of patient intake information associated with the user. For example, in some implementations, user video interfacing module 135 may be configured to monitor, on a frame-by-frame basis, the communication session via a videobot instance based on video and/or audio (e.g., corresponding to verbal communications and/or nonverbal communications) associated with a user at a user device (e.g., user device 110) in the communication session to detect interaction or gesture events for communication to patient intake device 120 and response (e.g., via query) to the events via videobot instruction generated by patient intake device 120” such that here for example AR options correspond to the algorithms used to detect interaction or gesture events which causes the system to respond interactively to the real life capture of the user via the videobot and thus provides an AR experience and such algorithms correspond to AR options that enable different aspects of AR such as different types of event recognitions as detailed in Table I in columns 13 and 14), Abbas teaches that it is known to select AR related processing options corresponding to processing a number of video frames per second and determining a number of video frames per second for applying the AR option to the accessed video frames based on a maximum rate of processing capability of the computing device and a length of time it takes to render the AR option and applying the AR option to each video frame of a subset of the accessed video frames at a predefined rate based on the determined number of video frames per second for applying the AR option (see Abbas, column 7, lines 50-67 through column 8, lines 1-65 teaching “User video interfacing module 135 receives, from user device 110 during the time period, a live multimedia stream associated with the user. In some implementations, user video interfacing module 135 generates instructions for execution to detect a user gesture event in connection with an act by the user during the time period based on at least one video frame of the live multimedia stream received during the time period” and “user video interfacing module 135 may generate the instructions for execution in real-time to detect a user gesture event based on audio and/or video of the user at user device 110 during the communication session. In some implementations, user video interfacing module 135 may implement one or more natural language processing algorithms, pattern matching or recognition algorithms, and a neural network to enable, establish, facilitate, and provide real-time video processing and video-based user-interfacing via a videobot. In such implementations, the natural language processing algorithms and the pattern matching or recognition algorithms may include, for example, algorithms configured for head gesture detection, hand gesture detection, medical device OCR detection, prescription detection, and/or speech recognition” such that here the system receives video from an accessed video stream of the user and their appearance and movements to provide AR feedback through a chatbot responding to the real life capture of the user in the video stream, and “Frame processing module 136 generates instructions for execution to determine a first frame process window (“FPW”) value based on a frame rate (e.g., in frames per second) of the rendering of the videobot instance by user device 110 during the first time period. An FPW value may be defined as the amount of time available to process, by one or more algorithms, a single frame (e.g., video frame) in a multimedia stream, which may be determined based on the frame rate (e.g., in frames per second) of the multimedia stream associated with the videobot instance” and also “Frame process times (“FPT(s)”), of the one or more algorithms may be determined by measurement of the time required by each of the one or more algorithms to process the single frame. The FPT of each algorithm can vary based on hardware, network conditions, video frame quality, and the like (e.g., of user device 110, of network 102), by which each algorithm may be implemented” where also an “actual frame process window (“AFPW”) may be defined as the actual time taken to process the single frame in the multimedia stream by the one or more algorithms combined, which may be determined by summation of the FPT of each algorithm” and further “the instructions may be executed to determine that the first FPW value exceeds a predetermined threshold. The predetermined threshold may correspond, for example, to the amount of time available to process each frame in the multimedia stream so as to provide, facilitate, and maintain real-time video processing, in accordance with embodiments of the present disclosure. In other words, the predetermined threshold may correspond to the amount of time made available by a particular device or system (e.g., user device 110) to process the single frame in the multimedia stream, which may depend on the limitations, capabilities, and resources of the device or system. Generally, the AFPW value for the one or more algorithms should always be less than the FPW value as may be determined based on a frame rate (e.g., in frames per second) of the rendering of the videobot instance by user device 110” such that here a predetermined threshold is then used to determine a number of video frames per second for applying the AR option and the AR option is then applied to each video frame of a subset of the accessed frames in a second time period that is based on the determined number of video frames per second for applying the AR option which will result in the option being skipped with regard to certain frames or the option being de-selected for a certain number of frames and this determination is based on a maximum processing capability of the system and the time it takes to render an AR option as this corresponds to “predetermined threshold may correspond to the amount of time made available by a particular device or system (e.g., user device 110) to process the single frame in the multimedia stream, which may depend on the limitations, capabilities, and resources of the device or system”; see further Abbas, column 12, lines 8-65 through column 13, lines 1-20 teaching “During a live video call, it is possible that the FPT of these algorithms are causing the actual frame process windows (“AFPW”) to exceed the FPW, which can cause unpredictable/unreliable results for event detection. This can also cause major performance concerns because the high AFPW can cause a backup of frames that need to be processed with no time to catch up” and “in response to determining that the first FPW value exceeds a predetermined threshold, then at 314, user device 110 performs a remedial action to define a frame rate of rendering the videobot instance during a second time period of the communication session after the first time period, such that a second FPW value of the videobot instance during the second time period, being based on the frame rate of the rendering the videobot instance during the second time period, is less than the first FPW value” and “predetermined threshold may be defined based on the first FPW value to thereby limit or maintain a magnitude of the first FPW value below that of the second FPW” and this remedial action performed in response to this determining “is performed to skip at least one frame during the rendering of the videobot instance during the second time period. For example, one or more orphan frames of the multimedia stream during the second time period may be skipped based on the delta between FPW and AFPW for event processing. For example, the remedial action can be performed to reduce the latency by reducing a video frame pixel density during the second time period below that of the first time period to help reduce the FPT for all algorithms” and “low priority algorithms may include non-required video processing processes, like certain gesture detections or object detections, which will not impact the quality of the call if they are disabled. For example, number recognition from a user's hands/fingers can be disabled if that is not used often” and as in column 14, lines 13-67 teaching “the remedial action is performed to skip execution of at least one image processing algorithm during the second time period based on a use value of the image processing algorithm relative to a use value of at least one other image processing algorithm” where “skipping execution of at least one image processing algorithm based on utility can reduce the AFPW time of each algorithm. Accordingly, the AFPW may be maintained under the FPW with respect to processing by a particular algorithm by disabling certain algorithms from processing” and “remedial action is performed in response to determining that an actual FPW is greater than a determined FPW. For example, if there are no more LPVAs to disable, then processing frames can be skipped based on the AFPW value. This will reduce the quality of the algorithms but will keep the processing of the video in real-time” such that here this skipping of frames or disabling of various processing options for the frames determines and applies the AR option to a subset of the accessed video frame based on the rate at which the option can be applied that will maintain the actual FPW below the threshold in the next time period, where again then this determination to skip a certain number of frames when applying the processing to the accessed video frame determines a number of video frames per second for applying the AR option and is based on the maximum rate of processing capability of the computing device running during the first time period and the length of time it takes to render AR options during the first time period such that during the second time period the AR option would only be applied to a subset of frames that are not skipped for that option). Thus Abbas teaches a known technique applicable to the base system of Du.
Therefore it would have been obvious for one of ordinary skill in the art before the effective filing date of the invention to modify Du by applying the known techniques of Abbas as doing so would be no more than application of a known technique to a base device ready for improvement which would yield predictable results and result in an improved system. The predictable results of applying Abbas’ technique to Du would be that the AR options of Du could be processed in relation to the real time video frames being captured to apply selected AR options as in Abbas where the AR options would be analyzed similarly to the AR based frame processing in Abbas such that the determining of the number of frames per second for applying the AR option and applying the AR option to a subset of accessed video frames would be applied as Abbas applies such AR processing at a determined number of video frames per second. Thus AR options would be selected to be applied to a subset of the image frames at a lower rate where the displayed result of the AR processing would be for each frame of the subset of video frames instead of from processing at a rate that would result in exceeding the maximum capabilities for processing of the system in relation to the amount of time it takes to render an AR option. This would result in an improved system as the system would be able to deal with changes in the frame rate that can occur from the environment and device and could more efficiently provide processing and could ensure that any AR options selected would be able to be applied in an amount of time that is actual available for processing as suggested by Abbas (see Abbas, column 14, lines 22-67 through column 15, lines 1-10 teaching “skipping execution of at least one image processing algorithm based on utility can reduce the AFPW time of each algorithm. Accordingly, the AFPW may be maintained under the FPW with respect to processing by a particular algorithm by disabling certain algorithms from processing” and “before the next frame arrives in the video stream, all the algorithms are to have already completed. Advantageously, this overcomes various of the aforementioned known problems, as buffering or re-sampling currently cannot be used in instant machine vision tasks to provide instant feedback to the user. This is due to this tendency that known algorithms and image processing libraries cannot be used. All processing by each detection algorithm are to be done frame-by-frame as they come” and “during a real-time communication session, video streams perform adaptive techniques to respond to network issues such as changing video quality and reducing frame rates. Because of this, the number of frames being received per second to the videobot constantly changes during the call, which can be responded to to maintain real-time video processing, in accordance with embodiments of the present disclosure. That is, if the frames are fewer, more time exists to process each frame. Also, if the frame rate increase, less time exists to process each frame”).
Regarding claim 4, Du as modified teaches all that is required as applied to claim 1 above and further teaches wherein the selection of the virtual camera is made via a list of camera devices in the user interface (see Du, paragraph 0018 teaching “generating an in-browser camera list that is selectable by a user. The in-browser camera list may include a list of cameras, for example a list of physical cameras internal or external to a computing device, received from an operating system kernel to which the virtual camera has been added. Upon receiving an indication that a browser is accessing a virtual camera from the in-browser camera list, the method may include receiving a physical camera frame from a physical camera associated with the virtual camera”).
Regarding claim 5, Du as modified teaches all that is required as applied to claim 1 above and further teaches wherein the selection of the AR option is made via a list of AR options displayed in the user interface (see Du, paragraph 0018 teaching the AR options which are the “local browser setting” which “may be selectable by a user and enable features such as filters, animations, visual text interpretations, and so forth” and as in paragraphs 0045-0048 the options are selectable by a user on the interface where “rowser application 254 may include a local browser setting 258. Local browser setting 258 may include one or more configurations relating to the virtual camera 262” and “local browser setting 258 relating to an animation may comprise a drop-down box selecting an animal or character” and “Sketch filter settings 112 in virtual camera setup 102 is one example of local browser setting 258” where selection from a drop-down list is selection via list and as in paragraph 0078 “the local browser setting is selectable in the browser and includes one or more of: a filter effect, an animation effect, a text effect, or a sound effect” such that here selection in the browser from such selectable options means they are displayed in some list form in order to select the listed option and such selection from a displayed list can also be seen in paragraph 0025 and figures 1A and 1B teaching “virtual camera setup 102 includes a filter selector 110, for which a sketch type filter is selected. As may be seen, the sketch filter has been applied to physical camera display 104 in virtual camera display 106, generating a modified frame that looks much like a sketch drawing. Other types of browser settings are possible as well, as further described below”).
Regarding claim 6, Du as modified teaches all that is required as applied to claim 1 above and further teaches wherein the user interface is displayed via the internet browser (see Du, paragraph 0078 teaching “the local browser setting is selectable in the browser and includes one or more of: a filter effect, an animation effect, a text effect, or a sound effect” and see paragraph 0025 and figures 1A and 1B teaching “local browser settings may facilitate the application of any combination of filter effects, animation effects, text effects, and/or sound effects. In the example, virtual camera setup 102 includes a filter selector 110, for which a sketch type filter is selected. As may be seen, the sketch filter has been applied to physical camera display 104 in virtual camera display 106, generating a modified frame that looks much like a sketch drawing. Other types of browser settings are possible as well, as further described below” such that it can be seen that the user interface is displayed via the internet browser where the program is being used).
Regarding claim 7, Du as modified teaches all that is required as applied to claim 1 above and further teaches generating an off-screen canvas element (note that this concept has already been addressed where the off-screen canvas element may be considered the “modified image” frame target before the frame is sent by the virtual camera output for actual display, recording, or transmission and thus this canvas element may be considered an off-screen canvas element as in paragraphs 0018-0020 teaching ”Next the modified frame may be sent displayed in the browser display and/or sent to another user participating in the video conference via another client application. The result is that the modified frames from the virtual camera are seamlessly substituted for the physical camera feed” and “virtual camera receives frames from a physical camera, modifies those physical camera frames, and outputs the modified frames. The modified frames may be displayed, saved, or sent to another computing device” such that here the “outputs the modified frames” refers to an off-screen canvas element in order to be sent to the new canvas element of the actual display target for the modified frame such as created by the display circuitry as in paragraph 0023 showing the “display of the modified frames” in the new canvas element “virtual camera display 106” as also taught in paragraphs 0062-0063 teaching “virtual camera 262 may render modified camera frame 308” and “may send modified camera frame 308 to a display processing module. The display processing module may comprise a browser display such as virtual camera display” such that again “modified camera frame 308” is rendered as an off-screen canvas element as it is not displayed until it is sent to the new canvas element which is the “browser display such as virtual camera display 106”); and
wherein the AR option is applied to each video frame of the subset of the accessed video frames in the off-screen canvas element and provided from the off-screen canvas element to the new canvas element (see Du, paragraphs 0018-0020 teaching ”Next the modified frame may be sent displayed in the browser display and/or sent to another user participating in the video conference via another client application. The result is that the modified frames from the virtual camera are seamlessly substituted for the physical camera feed” and “virtual camera receives frames from a physical camera, modifies those physical camera frames, and outputs the modified frames. The modified frames may be displayed, saved, or sent to another computing device” such that here the “outputs the modified frames” refers to an off-screen canvas element in order to be sent to the new canvas element of the actual display target for the modified frame such as created by the display circuitry as in paragraph 0023 showing the “display of the modified frames” in the new canvas element “virtual camera display 106” as also taught in paragraphs 0062-0063 teaching “virtual camera 262 may render modified camera frame 308” and “may send modified camera frame 308 to a display processing module. The display processing module may comprise a browser display such as virtual camera display” such that again “modified camera frame 308” is rendered as an off-screen canvas element as it is not displayed until it is sent to the new canvas element which is the “browser display such as virtual camera display 106”).
Regarding claim 8, Du as modified teaches all that is required as applied to claim 1 above and further teaches wherein the AR option is a first AR option (see Du, paragraphs 0018-0020 teaching “modified frame may be generated based on the physical camera frame and a local browser setting. The local browser setting may be selectable by a user and enable features such as filters, animations, visual text interpretations, and so forth” where the local browser settings are AR options and a user may select one of them as in paragraphs 0022-0024 teaching an AR option selected as a “sketch effect”)and the operations further comprise:
detecting a selection of a second AR option (see Du, paragraphs 0022-0030 and figures 1A and 1B where a second AR option has been selected as where “Virtual camera display 107 differs from virtual camera display 106 in that it includes a text effect 116. Further text effect 116 includes a text transliteration of speech from the first user. This may be further seen by the filter selected in filter selector 111, which is “sketch with text.” In examples, it may be possible to combine any number of effects in virtual camera display 107”);
applying the second AR option to each video frame of at least a second subset of the accessed video frames, at the predefined rate, based on the determined number of video frames per second for applying the AR option (see Du, paragraphs 0022-0030 and figures 1A and 1B where a second AR option has been applied to each video frame of the second subset of frames that is now being accessed from the physical camera and applied at a predefined rate corresponding to the frames per second of the video feed where “Virtual camera display 107 differs from virtual camera display 106 in that it includes a text effect 116. Further text effect 116 includes a text transliteration of speech from the first user. This may be further seen by the filter selected in filter selector 111, which is “sketch with text.” In examples, it may be possible to combine any number of effects in virtual camera display 107” and as already modified by Abbas above, such second option is already within the combination of Du and Abbas as Abbas teaches various algorithms can be applied to a subset of frames and application of the AR option would be at the predefined rate in any situation in which the options do not exceed the predetermined threshold as for example where Abbas teaches “the remedial action is performed to skip execution of at least one image processing algorithm during the second time period based on a use value of the image processing algorithm relative to a use value of at least one other image processing algorithm. For example, the use value may be used as a measure of relative utility or performance of each of the algorithms, such as with respect to detecting a particular interaction or gesture event. In other words, skipping execution of at least one image processing algorithm based on utility can reduce the AFPW time of each algorithm. Accordingly, the AFPW may be maintained under the FPW with respect to processing by a particular algorithm by disabling certain algorithms from processing. We determine the priority of each algorithm based on certain factors, then enable/disable algorithms as needed during the video call, until the AFPW is below the FPW”);
providing each video frame of the second subset of the accessed video frames comprising the applied second AR option to the new canvas element (see Du, paragraphs 0022-0028 and figures 1A and 1B above teaching “Virtual camera display 107 differs from virtual camera display 106 in that it includes a text effect 116. Further text effect 116 includes a text transliteration of speech from the first user. This may be further seen by the filter selected in filter selector 111, which is “sketch with text.” In examples, it may be possible to combine any number of effects in virtual camera display 107” such that each video frame now has the second AR option applied and this is provided to the new canvas element of the virtual camera modified frame canvas element to apply the effect and is also sent to the new canvas element of the display target as explained above in paragraphs 0062-0063 teaching “virtual camera 262 may render modified camera frame 308” and “may send modified camera frame 308 to a display processing module. The display processing module may comprise a browser display such as virtual camera display” such that again “modified camera frame 308” is rendered as a new canvas element and furthermore these frames as modified are sent to the new canvas element of the rendered display such as the actual browser display on the user interface); and
causing display, of each video frame of the second subset of the accessed video frames comprising the applied second AR option in the new canvas element, on the user interface of the computing device (see Du, paragraphs 0022-0028 and figures 1A and 1B above teaching “Virtual camera display 107 differs from virtual camera display 106 in that it includes a text effect 116. Further text effect 116 includes a text transliteration of speech from the first user. This may be further seen by the filter selected in filter selector 111, which is “sketch with text.” In examples, it may be possible to combine any number of effects in virtual camera display 107” such that each video frame now has the second AR option applied and this is provided to the new canvas element of the virtual camera modified frame canvas element to apply the effect and is also sent to the new canvas element of the display target as explained above in paragraphs 0062-0063 teaching “virtual camera 262 may render modified camera frame 308” and “may send modified camera frame 308 to a display processing module. The display processing module may comprise a browser display such as virtual camera display” such that again “modified camera frame 308” is rendered as a new canvas element and furthermore these frames as modified are sent to the new canvas element of the rendered display such as the actual browser display on the user interface).
Regarding claim 9, Du as modified teaches all that is required as applied to claim 1 above and further teaches wherein the selection of the virtual camera is detected based on identifying a camera device identifier associated with the virtual camera (see Du, paragraphs 0050-0054 teaching “virtual camera initialization module” which “may be executed prior to virtual camera use in a web application to initialize in-browser camera list 256. Virtual camera initialization module 259 generates in-browser camera list 256 and adds virtual camera 262 to in-browser camera list 256” and “virtual camera initialization module 259 may receive a camera list from operating system kernel 250, for example operating system camera list 252. Virtual camera initialization module 259 may add the operating system camera list 252 to the in-browser camera list” and “Virtual camera initialization module 259 may be executed on startup of browser application 254 or upon initialization of a web application that uses a camera” such that here the virtual camera is added to a list for selection so that it can be identified as a camera device which allows detection of selection from a list identifying the cameras including the virtual camera which is available for use by the browser as in paragraph 0044 and figure 1A teaching “Browser application 254 may include an in-browser camera list 256. In examples, the in-browser camera list 256 may include a combination of operating system camera list 252 and virtual camera 262. In-browser camera list 256 may be selectable by a user within browser application 254. For example, FIG. 1A depicts in-browser camera selectable list 108, which may be used to select a camera from in-browser camera list 256”).
Regarding claim 10, Du as modified teaches all that is required as applied to claim 1 above and further teaches wherein the hardware camera coupled with the computing device is a hardware camera last used by a user of the computing device (note that the claim does not define what the “use” of the last hardware camera was or when “last” refers to such that if any point in the past if a hardware camera was used by the computing device and that hardware camera is coupled with the computing device then the limitation is met; note that the claim does not require any specific determination that the device is a hardware camera in some way last used by a user of the computing device and merely requires a coincidental or common situation and does not limit or define when “last” is with regard to and furthermore does not tie such limitation back to any functioning of the claim language; thus see Du, paragraph 0054 teaching “the physical camera that is the default camera in operating system camera list 252 or in-browser camera list 256 may be associated with virtual camera 262 via local browser camera setting 260 at initialization of browser application 254” and as in paragraph 0018 “Upon receiving an indication that a browser is accessing a virtual camera from the in-browser camera list, the method may include receiving a physical camera frame from a physical camera associated with the virtual camera” which means that a default hardware camera coupled with the computing device is the hardware camera and as it is the default camera this means that at least it is the last hardware camera used by the user of the computing device as the default camera; furthermore note that as the system is capable of being used multiple times then at any time after the default camera has been used for the process at least once, then when running the system again this would make that default camera the last hardware camera coupled that was used by the computing device for the more specific purpose of the processing in relation to the virtual camera).
Regarding claims 11-12, and 14-19, the instant claims correspond to a “computer-implemented method comprising” a series of functions to be implemented by a computer where the functions performed correspond to the function performed by the system of claims 1,2, 4-8 and 10, respectively. Du as modified teaches such a computer-implemented method as implemented by the computer system as already addressed in the rejections of claims 1,2, 4-8 and 10 above. In light of this, the limitations of claims 11-12, and 14-19 correspond to the limitations of claims 1,2, 4-8 and 10, respectively; thus they are rejected on the same grounds as claims 1,2, 4-8 and 10, respectively.
Regarding claims 20-24, the instant claims are recited as an apparatus in the form of a “non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising” the operations as addressed in the rejection of the computer system embodiment of claim 1. Du as modified teaches such a non-transitory CRM as already addressed above (see Du, paragraph 0034 “processor 204 may include multiple processors, and memory 206 may include multiple memories. Processor 204 may be in communication with any cameras, sensors, and other modules and electronics of computing device 202. Processor 204 is configured by instructions (e.g., software, application, modules, etc.) to execute a virtual camera. The instructions may include non-transitory computer readable instructions stored in, and recalled from, memory 206” and paragraph 0090 teaching “Note also that the software implemented aspects of the example implementations are typically encoded on some form of non-transitory program storage medium or implemented over some type of transmission medium. The program storage medium may be magnetic (e.g., a floppy disk or a hard drive) or optical (e.g., a compact disk read only memory, or CD ROM), and may be read only or random access. Similarly, the transmission medium may be twisted wire pairs, coaxial cable, optical fiber, or some other suitable transmission medium known to the art. The example implementations not limited by these aspects of any given implementation”). In light of this, the limitations of claims 20 and 21-24 correspond to the limitations of claims 1, and 7-10, respectively; thus they are rejected on the same grounds as claims 1, and 7-10, respectively.
Response to Arguments
Applicant’s arguments, see “REMARKS”, filed 4/16/2026, with respect to the rejection(s) of claim(s) 1-2, 4-12, and 14-20 under 35 U.S.C. 102 as being anticipated by Du have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Du and Abbas as explained above.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. As noted previously, see Swierk et al (US Patent NO. 11350059, teaching a videoconferencing system in which captured image frames are subjected to AR like effect processing and a computational burden of the effects to be applied can be used to adjust the frame rate setting that the effects are applied with). See also Yang et al (US PGPUB No. 20240320932) (see paragraphs 0048-0071 updating a frame rate at which augmented reality processing is applied to incoming frames based on keeping frames within a target frame rate based on rendering being applied).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SCOTT E SONNERS whose telephone number is (571)270-7504. The examiner can normally be reached Mon-Friday 9-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached at (571) 272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SCOTT E SONNERS/Examiner, Art Unit 2613
/XIAO M WU/Supervisory Patent Examiner, Art Unit 2613
1 WO 2024/177668, corresponding to Foreign Patent Document 1 in the IDS filed 8/22/2025
2 US Patent No. 11315692