DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
The objections to the drawings and the specifications have been withdrawn in view of the applicants amendments filed on 07/01/2026.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 2, 5-6, 8, 10-11, 14-16, 18-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Molholm (U.S. Pub. No 20210099632) in view of Ollila (U.S. Pub. No. 20220343529) and Hirvonen et al. (U.S. Pub. No 20250124667).
Regarding claim 1, Molholm discloses a system comprising: at least one physical processor; and physical memory comprising computer-executable instructions that, when executed by the physical processor, cause the physical processor to (paragraph 62, “In some embodiments, HMD 2000 may include a controller 2030 configured to implement functionality of the MR system and to generate frames (each frame including a left and right image) that are provided to displays 2022A and 2022B. In some embodiments, HMD 2000 may also include a memory 2032 configured to store software (code 2034) of the MR system that is executable by the controller 2030, as well as data 2038 that may be used by the MR system when executing on the controller 2030.”; also, para 63, “In various embodiments, controller 2030 may be a uniprocessor system including one processor, or a multiprocessor system including several processors (e.g., two, four, eight, or another suitable number).”; also, para 63, “Controller 2030 may include one or more processing cores each configured to execute instructions”): receive a frame of a passthrough video captured by a sensor of a mixed-reality display (paragraph 21, “Embodiments may, for example, be implemented in MR systems that include a head mounted display (HMD) equipped with scene cameras for video pass-through, an eye or gaze tracking system, and a method for ambient light detection such as one or more ambient light sensors.”; also, para 45, “In some embodiments, video streams of the real environment captured by the visible light cameras 150 may be processed by the controller 160 of the HMD 100 to render augmented or mixed reality frames that include virtual content overlaid on the view of the real environment, and the rendered frames may be provided to display 110.”); frame by at least one visual adjustment algorithm (para 47, “In the display pipeline 280, exposure compensation 282 is applied to the image from the camera 250 (after ISP 262 processing without tone mapping) to scale the image to the proper scene exposure. Exposure compensation 282 is performed with adequate precision to be lossless to the image.”) adjusted frame on the mixed-reality display (paragraph 55, “As indicated at 330, the blended image is displayed.”). Molholm does not disclose detect at least one region of the frame of the passthrough video that is occluded by a virtual object projected within the mixed-reality display; create a modified version of the frame of the passthrough video without the region that is occluded by the virtual object; provide the modified version of the frame to at least one visual analysis algorithm that calculates at least one statistic based on the modified version of the frame, based at least in part on the at least one statistic calculated based on the modified version.
However, in a similar field of endeavor, Ollila discloses detect at least one region of the frame of the passthrough video that is occluded by a virtual object projected within the mixed-reality display (para 49, “Throughout the present disclosure, the term "image segment" refers to a part of the at least one image over which the blend object is to be superimposed.”; also, para 59, “In other implementations, the blend object to be superimposed is an opaque object. Upon superimposition of such a blend object over the image segment, the image segment would be fully obscured and a portion of the real-world scene corresponding to the image segment would not be visible in the at least one image.,”; also, para 50, “The at least one region within the photo-sensitive surface that corresponds to the image segment is determined by mapping the location of the image segment in the at least one image to corresponding pixels in the photo-sensitive surface.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Molholm's invention of a mixed reality system having a physical processor and a memory storing executable code, which receives a frame of a passthrough video captured by a scene camera of a head mounted display, adjusts the frame by an exposure compensation applied in the display pipeline, and displays the adjusted frame, with the features of Ollila's invention of determining a part of the image over which a blend object is to be superimposed and which, where that blend object is opaque, that blend object fully obscures. One of ordinary skill would have been led to make this modification because Molholm's stated purpose is to acquire a properly exposed image of the real world scene, and Ollila teaches that where the superimposed object is opaque the corresponding portion of the real-world scene would not be visible in the resulting image and that the pixels of such a part are for that reason treated differently from the pixels of the remaining region in the imaging chain, so that identifying that part of Molholm's frame serves Molholm's own purpose of properly exposing the scene rather than defeating it. The results of the modification would have been predictable to one of ordinary skill, who would have had a reasonable expectation of success, because Molholm itself renders the virtual content it composites at blend 284 and therefore already has available the location and the extent of that content on the frame, so that the part Ollila determines can be obtained in Molholm from data Molholm already generates and requires no additional sensing and no additional hardware, and because Molholm's gaze-positioned region of interest, its camera auto-exposure, its exposure compensation, and its blending operation each continue to operate as Molholm describes.
Hirvonen discloses create a modified version of the frame of the passthrough video without the region that is occluded by the virtual object (para 108, “It will also be appreciated that the previously-captured image may correspond to a different viewpoint as compared to the given viewpoint; therefore, the previously-captured image needs to be reprojected accordingly, in order to match a perspective of the different viewpoint with the perspective of the given viewpoint, for the at least one processor to determine the image segment of the previously-captured image that corresponds to the remaining portion of the field of view.”; also, para 102, “In addition to this, since some real objects or their portions would highly likely be visible from the perspective of the given viewpoint (and would be represented in the image) as they not being occluded by the set of virtual objects, the at least one processor is configured to read out pixel data from the remaining region of the photo-sensitive surface.”); provide the modified version of the frame to at least one visual analysis algorithm that calculates at least one statistic based on the modified version of the frame (para 105, “Alternatively or additionally, optionally, when capturing the image by controlling the at least one parameter at (ii), the at least one processor is configured to adjust at least one of: exposure time, sensitivity, aperture size, based on the saturation of the pixels in the image segment of the previously-captured image that corresponds to the remaining portion of the field of view.”), based at least in part on the at least one statistic calculated based on the modified version (para 109, “Yet alternatively or additionally, optionally, when capturing the image by controlling the at least one parameter at (ii), the at least one processor is configured to adjust white balance, based on the colour temperature of the pixels in the image segment of the previously-captured image that corresponds to the remaining portion of the field of view.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Molholm in view of Ollila, in which a mixed reality system receives a frame of passthrough video and identifies the part of that frame over which superimposed content fully obscures the real world scene, with the features of Hirvonen's invention of determining the image segment that corresponds to the remaining portion of the field of view, reading out pixel data from the remaining region, and deriving the saturation and the colour temperature of the pixels in that segment to set exposure and white balance. The combination would have been obvious because the base combination identifies which part of the frame is obscured but still leaves that obscured part in the input to the camera control computation, and Hirvonen supplies precisely the step the base combination lacks by confining the derivation of the saturation and colour temperature values to the segment corresponding to the portion of the field of view that is not occluded. A person of ordinary skill would have recognized that pixels lying under superimposed content carry the values of that content rather than of the physical scene, so that confining the derivation to the remaining segment yields a control value that reflects the real world scene, with the predictable result of a more accurate exposure and a more accurate colour reproduction in the frame that Molholm then displays. That person would have had a reasonable expectation of success, because Hirvonen derives the saturation and the colour temperature by the same ordinary measurements Molholm already performs on its region of interest, and the base combination has already identified the part of the frame to be left out of them.
Regarding claim 2, Molholm as modified by Ollila and Hirvonen discloses the system of claim 1, wherein Molholm further discloses the at least one visual adjustment algorithm comprises an automatic exposure control algorithm that adjusts an exposure level of the frame (para 47, “In the display pipeline 280, exposure compensation 282 is applied to the image from the camera 250 (after ISP 262 processing without tone mapping) to scale the image to the proper scene exposure. Exposure compensation 282 is performed with adequate precision to be lossless to the image.”; also, para 47, “In the image output by exposure compensation 282, the region of interest in the scene remains as auto-exposed by the camera, while the rest of the image outside the region of interest is compensated to an exposure (referred to as scene exposure) as determined form the ambient light information.”).
Regarding claim 5, Molholm as modified by Ollila and Hirvonen discloses the system of claim 1, wherein Molholm further discloses the adjusted frame comprises the virtual object (Molholm: paragraph 48, “In the rendering pipeline 270, virtual content 271 may be rendered into an image to be blended with the image captured by the camera 250 in the display pipeline 280. Exposure compensation 272 is applied so that the rendered virtual content has the same scene exposure as the exposure-compensated image in the display pipeline 280.”; also, paragraph 49, “In the display pipeline 280, the rendered virtual content is blended 284 into the exposure-compensated image, for example using an additive alpha blend (Aa+B(1−a)).”).
Regarding claim 6, Molholm as modified by Ollila and Hirvonen discloses the system of claim 1, wherein Molholm further discloses: the mixed-reality display comprises a head-mounted display; and the passthrough video comprises video of a physical environment captured by at least one sensor of the head-mounted display (paragraph 20, “Embodiments may, for example, be implemented in MR systems that include a head mounted display (HMD) equipped with scene cameras for video pass-through, an eye or gaze tracking system, and a method for ambient light detection such as one or more ambient light sensors.”; also, para 65, “In some embodiments, the HMD 2000 may include one or more sensors 2050 that collect information about the user's environment (video, depth information, lighting information, etc.). The sensors 2050 may provide the information to the controller 2030 of the MR system. In some embodiments, sensors 2050 may include, but are not limited to, visible light cameras (e.g., video cameras) and ambient light sensors.”).
Regarding claim 8, Molholm discloses a method comprising (para 4, “Embodiments of a processing pipeline and method for MR systems that utilizes selective auto-exposure for a region of interest in a scene based on gaze and that compensates exposure for the rest of the scene based on ambient lighting information for the scene are described.”): receiving a frame of a passthrough video captured by a sensor of a mixed-reality display (paragraph 20, “Embodiments may, for example, be implemented in MR systems that include a head mounted display (HMD) equipped with scene cameras for video pass-through, an eye or gaze tracking system, and a method for ambient light detection such as one or more ambient light sensors.”; also, para 45, “In some embodiments, video streams of the real environment captured by the visible light cameras 150 may be processed by the controller 160 of the HMD 100 to render augmented or mixed reality frames that include virtual content overlaid on the view of the real environment, and the rendered frames may be provided to display 110.”); frame by at least one visual adjustment algorithm (paragraph 47, “In the display pipeline 280, exposure compensation 282 is applied to the image from the camera 250 (after ISP 262 processing without tone mapping) to scale the image to the proper scene exposure.”; also, paragraph 47, “In the image output by exposure compensation 282, the region of interest in the scene remains as auto-exposed by the camera, while the rest of the image outside the region of interest is compensated to an exposure (referred to as scene exposure) as determined form the ambient light information.”) adjusted frame on the mixed-reality display (paragraph 55, “As indicated at 330, the blended image is displayed.”). Molholm does not disclose detecting at least one region of the frame of the passthrough video that is occluded by a virtual object projected within the mixed-reality display; creating a modified version of the frame of the passthrough video without the region that is occluded by the virtual object; providing the modified version of the frame to at least one visual analysis algorithm that calculates at least one statistic based on the modified version of the frame; based at least in part on the at least one statistic calculated based on the modified version of the frame.
However, in a similar field of endeavor, Ollila discloses detecting at least one region of the frame of the passthrough video that is occluded by a virtual object projected within the mixed-reality display (para 49, “Throughout the present disclosure, the term "image segment" refers to a part of the at least one image over which the blend object is to be superimposed.”; also, para 59, “In other implementations, the blend object to be superimposed is an opaque object. Upon superimposition of such a blend object over the image segment, the image segment would be fully obscured and a portion of the real-world scene corresponding to the image segment would not be visible in the at least one image.,”; also, para 50, “The at least one region within the photo-sensitive surface that corresponds to the image segment is determined by mapping the location of the image segment in the at least one image to corresponding pixels in the photo-sensitive surface.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Molholm's invention of a mixed reality method that receives a frame of a passthrough video captured by a scene camera of a head mounted display, adjusts the frame by an exposure compensation applied in the display pipeline, and displays the adjusted frame, with the features of Ollila's invention of determining a part of the image over which a blend object is to be superimposed and which, where that blend object is opaque, that blend object fully obscures. One of ordinary skill would have been led to make this modification because Molholm's stated purpose is to acquire a properly exposed image of the real world scene, and Ollila teaches that where the superimposed object is opaque the corresponding portion of the real-world scene would not be visible in the resulting image and that the pixels of such a part are for that reason treated differently from the pixels of the remaining region in the imaging chain, so that identifying that part of Molholm's frame serves Molholm's own purpose of properly exposing the scene rather than defeating it. The results of the modification would have been predictable to one of ordinary skill, who would have had a reasonable expectation of success, because Molholm itself renders the virtual content it composites at blend 284 and therefore already has available the location and the extent of that content on the frame, so that the part Ollila determines can be obtained in Molholm from data Molholm already generates and requires no additional sensing and no additional hardware, and because Molholm's gaze-positioned region of interest, its camera auto-exposure, its exposure compensation, and its blending operation each continue to operate as Molholm describes.
Hirvonen discloses creating a modified version of the frame of the passthrough video without the region that is occluded by the virtual object (para 108, “It will also be appreciated that the previously-captured image may correspond to a different viewpoint as compared to the given viewpoint; therefore, the previously-captured image needs to be reprojected accordingly, in order to match a perspective of the different viewpoint with the perspective of the given viewpoint, for the at least one processor to determine the image segment of the previously-captured image that corresponds to the remaining portion of the field of view.”; also, para 102, “In addition to this, since some real objects or their portions would highly likely be visible from the perspective of the given viewpoint (and would be represented in the image) as they not being occluded by the set of virtual objects, the at least one processor is configured to read out pixel data from the remaining region of the photo-sensitive surface.”); providing the modified version of the frame to at least one visual analysis algorithm that calculates at least one statistic based on the modified version of the frame (para 105, “Alternatively or additionally, optionally, when capturing the image by controlling the at least one parameter at (ii), the at least one processor is configured to adjust at least one of: exposure time, sensitivity, aperture size, based on the saturation of the pixels in the image segment of the previously-captured image that corresponds to the remaining portion of the field of view.”); based at least in part on the at least one statistic calculated based on the modified version of the frame (para 109, “Yet alternatively or additionally, optionally, when capturing the image by controlling the at least one parameter at (ii), the at least one processor is configured to adjust white balance, based on the colour temperature of the pixels in the image segment of the previously-captured image that corresponds to the remaining portion of the field of view.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Molholm in view of Ollila, in which a mixed reality method receives a frame of passthrough video and identifies the part of that frame over which superimposed content fully obscures the real world scene, with the features of Hirvonen's invention of determining the image segment that corresponds to the remaining portion of the field of view, reading out pixel data from the remaining region, and deriving the saturation and the colour temperature of the pixels in that segment to set exposure and white balance. The combination would have been obvious because the base combination identifies which part of the frame is obscured but still leaves that obscured part in the input to the camera control computation, and Hirvonen supplies precisely the step the base combination lacks by confining the derivation of the saturation and colour temperature values to the segment corresponding to the portion of the field of view that is not occluded. A person of ordinary skill would have recognized that pixels lying under superimposed content carry the values of that content rather than of the physical scene, so that confining the derivation to the remaining segment yields a control value that reflects the real world scene, with the predictable result of a more accurate exposure and a more accurate colour reproduction in the frame that Molholm then displays. That person would have had a reasonable expectation of success, because Hirvonen derives the saturation and the colour temperature by the same ordinary measurements Molholm already performs on its region of interest, and the base combination has already identified the part of the frame to be left out of them.
Regarding claim 10, Molholm as modified by Ollila and Hirvonen discloses the method of claim 8, wherein Molholm further discloses the virtual object comprises a virtual environment (para 2, “Virtual reality (VR) allows users to experience and/or interact with an immersive artificial environment, such that the user feels as if they were physically in that environment.”; also, para 69, “Embodiments of the HMD 2000 as illustrated in FIG. 5 may also be used in virtual reality (VR) applications to provide VR views to the user. In these embodiments, the controller 2030 of the HMD 2000 may render or obtain virtual reality (VR) frames that include virtual content, and the rendered frames may be displayed to provide a virtual reality (as opposed to mixed reality) experience to the user.”).
Regarding claim 11, Molholm as modified by Ollila and Hirvonen discloses the method of claim 8, wherein Molholm further discloses the at least one visual adjustment algorithm comprises an automatic exposure control algorithm that adjusts an exposure level of the frame (para 47, “In the display pipeline 280, exposure compensation 282 is applied to the image from the camera 250 (after ISP 262 processing without tone mapping) to scale the image to the proper scene exposure. Exposure compensation 282 is performed with adequate precision to be lossless to the image.”; also, para 47, “In the image output by exposure compensation 282, the region of interest in the scene remains as auto-exposed by the camera, while the rest of the image outside the region of interest is compensated to an exposure (referred to as scene exposure) as determined form the ambient light information.”).
Regarding claim 14, Molholm as modified by Ollila and Hirvonen discloses the method of claim 8, wherein Molholm further discloses the adjusted frame comprises the virtual object (para 48, “In the rendering pipeline 270, virtual content 271 may be rendered into an image to be blended with the image captured by the camera 250 in the display pipeline 280. Exposure compensation 272 is applied so that the rendered virtual content has the same scene exposure as the exposure-compensated image in the display pipeline 280.”; also, para 49, “In the display pipeline 280, the rendered virtual content is blended 284 into the exposure-compensated image, for example using an additive alpha blend (Aa+B(1−a)).”).
Regarding claim 15, Molholm as modified by Ollila and Hirvonen discloses the method of claim 8, wherein Molholm further discloses: the mixed-reality display comprises a head-mounted display; and the passthrough video comprises video of a physical environment captured by at least one sensor of the head-mounted display (para 20, “Embodiments may, for example, be implemented in MR systems that include a head mounted display (HMD) equipped with scene cameras for video pass-through, an eye or gaze tracking system, and a method for ambient light detection such as one or more ambient light sensors.”; also, para 65, “In some embodiments, the HMD 2000 may include one or more sensors 2050 that collect information about the user's environment (video, depth information, lighting information, etc.). The sensors 2050 may provide the information to the controller 2030 of the MR system. In some embodiments, sensors 2050 may include, but are not limited to, visible light cameras (e.g., video cameras) and ambient light sensors.”).
Regarding claim 16, Molholm discloses a non-transitory computer-readable medium comprising one or more computer-readable instructions that, when executed by at least one processor of a computing device, cause the computing device to (para 62, “In some embodiments, HMD 2000 may also include a memory 2032 configured to store software (code 2034) of the MR system that is executable by the controller 2030, as well as data 2038 that may be used by the MR system when executing on the controller 2030.”; also, para 63, “In various embodiments, controller 2030 may be a uniprocessor system including one processor, or a multiprocessor system including several processors (e.g., two, four, eight, or another suitable number).”; also, para 63, “Controller 2030 may include one or more processing cores each configured to execute instructions.”; also, para 70, “Generally speaking, computer-readable media may include non-transitory, computer-readable storage media or memory media such as magnetic or optical media, e.g., disk or DVD/CD-ROM, volatile or non-volatile media such as RAM (e.g. SDRAM, DDR, RDRAM, SRAM, etc.), ROM, etc.”): receive a frame of a passthrough video captured by a sensor of a mixed-reality display (paragraph 20, “Embodiments may, for example, be implemented in MR systems that include a head mounted display (HMD) equipped with scene cameras for video pass-through, an eye or gaze tracking system, and a method for ambient light detection such as one or more ambient light sensors.”; also, para 45, “In some embodiments, video streams of the real environment captured by the visible light cameras 150 may be processed by the controller 160 of the HMD 100 to render augmented or mixed reality frames that include virtual content overlaid on the view of the real environment, and the rendered frames may be provided to display 110.”); frame by at least one visual adjustment algorithm (para 47, “In the display pipeline 280, exposure compensation 282 is applied to the image from the camera 250 (after ISP 262 processing without tone mapping) to scale the image to the proper scene exposure. Exposure compensation 282 is performed with adequate precision to be lossless to the image.”) adjusted frame on the mixed-reality display (para 55, “As indicated at 330, the blended image is displayed.”). Molholm does not disclose detect at least one region of the frame of the passthrough video that is occluded by a virtual object projected within the mixed-reality display; create a modified version of the frame of the passthrough video without the region that is occluded by the virtual object; provide the modified version of the frame to at least one visual analysis algorithm that calculates at least one statistic based on the modified version of the frame; based at least in part on the at least one statistic calculated based on the modified version.
However, in a similar field of endeavor, Ollila discloses detect at least one region of the frame of the passthrough video that is occluded by a virtual object projected within the mixed-reality display (para 49, “Throughout the present disclosure, the term "image segment" refers to a part of the at least one image over which the blend object is to be superimposed.”; also, para 59, “In other implementations, the blend object to be superimposed is an opaque object. Upon superimposition of such a blend object over the image segment, the image segment would be fully obscured and a portion of the real-world scene corresponding to the image segment would not be visible in the at least one image.,”; also, para 50, “The at least one region within the photo-sensitive surface that corresponds to the image segment is determined by mapping the location of the image segment in the at least one image to corresponding pixels in the photo-sensitive surface.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Molholm's invention of a memory storing executable code of a mixed reality system that receives a frame of a passthrough video captured by a scene camera of a head mounted display, adjusts the frame by an exposure compensation applied in the display pipeline, and displays the adjusted frame, with the features of Ollila's invention of determining a part of the image over which a blend object is to be superimposed and which, where that blend object is opaque, that blend object fully obscures. One of ordinary skill would have been led to make this modification because Molholm's stated purpose is to acquire a properly exposed image of the real world scene, and Ollila teaches that where the superimposed object is opaque the corresponding portion of the real-world scene would not be visible in the resulting image and that the pixels of such a part are for that reason treated differently from the pixels of the remaining region in the imaging chain, so that identifying that part of Molholm's frame serves Molholm's own purpose of properly exposing the scene rather than defeating it. The results of the modification would have been predictable to one of ordinary skill, who would have had a reasonable expectation of success, because Molholm itself renders the virtual content it composites at blend 284 and therefore already has available the location and the extent of that content on the frame, so that the part Ollila determines can be obtained in Molholm from data Molholm already generates and requires no additional sensing and no additional hardware, and because Molholm's gaze-positioned region of interest, its camera auto-exposure, its exposure compensation, and its blending operation each continue to operate as Molholm describes.
Hirvonen discloses create a modified version of the frame of the passthrough video without the region that is occluded by the virtual object (para 108, “It will also be appreciated that the previously-captured image may correspond to a different viewpoint as compared to the given viewpoint; therefore, the previously-captured image needs to be reprojected accordingly, in order to match a perspective of the different viewpoint with the perspective of the given viewpoint, for the at least one processor to determine the image segment of the previously-captured image that corresponds to the remaining portion of the field of view.”; also, para 102, “In addition to this, since some real objects or their portions would highly likely be visible from the perspective of the given viewpoint (and would be represented in the image) as they not being occluded by the set of virtual objects, the at least one processor is configured to read out pixel data from the remaining region of the photo-sensitive surface.”); provide the modified version of the frame to at least one visual analysis algorithm that calculates at least one statistic based on the modified version of the frame (para 105, “Alternatively or additionally, optionally, when capturing the image by controlling the at least one parameter at (ii), the at least one processor is configured to adjust at least one of: exposure time, sensitivity, aperture size, based on the saturation of the pixels in the image segment of the previously-captured image that corresponds to the remaining portion of the field of view.”); based at least in part on the at least one statistic calculated based on the modified version (para 109, “Yet alternatively or additionally, optionally, when capturing the image by controlling the at least one parameter at (ii), the at least one processor is configured to adjust white balance, based on the colour temperature of the pixels in the image segment of the previously-captured image that corresponds to the remaining portion of the field of view.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Molholm in view of Ollila, in which stored executable code causes a mixed reality system to receive a frame of passthrough video and identify the part of that frame over which superimposed content fully obscures the real world scene, with the features of Hirvonen's invention of determining the image segment that corresponds to the remaining portion of the field of view, reading out pixel data from the remaining region, and deriving the saturation and the colour temperature of the pixels in that segment to set exposure and white balance. The combination would have been obvious because the base combination identifies which part of the frame is obscured but still leaves that obscured part in the input to the camera control computation, and Hirvonen supplies precisely the step the base combination lacks by confining the derivation of the saturation and colour temperature values to the segment corresponding to the portion of the field of view that is not occluded. A person of ordinary skill would have recognized that pixels lying under superimposed content carry the values of that content rather than of the physical scene, so that confining the derivation to the remaining segment yields a control value that reflects the real world scene, with the predictable result of a more accurate exposure and a more accurate colour reproduction in the frame that Molholm then displays. That person would have had a reasonable expectation of success, because Hirvonen derives the saturation and the colour temperature by the same ordinary measurements Molholm already performs on its region of interest, and the base combination has already identified the part of the frame to be left out of them.
Regarding claim 18, Molholm as modified by Ollila and Hirvonen discloses the non-transitory computer-readable medium of claim 16, wherein Molholm further discloses the virtual object comprises a virtual environment (para 2, “Virtual reality (VR) allows users to experience and/or interact with an immersive artificial environment, such that the user feels as if they were physically in that environment.”; also, 69, “Embodiments of the HMD 2000 as illustrated in FIG. 5 may also be used in virtual reality (VR) applications to provide VR views to the user. In these embodiments, the controller 2030 of the HMD 2000 may render or obtain virtual reality (VR) frames that include virtual content, and the rendered frames may be displayed to provide a virtual reality (as opposed to mixed reality) experience to the user.”).
Regarding claim 19, Molholm as modified by Ollila and Hirvonen discloses the non-transitory computer-readable medium of claim 16, wherein Molholm further discloses the at least one visual adjustment algorithm comprises an automatic exposure control algorithm that adjusts an exposure of the frame (para 47, “In the display pipeline 280, exposure compensation 282 is applied to the image from the camera 250 (after ISP 262 processing without tone mapping) to scale the image to the proper scene exposure. Exposure compensation 282 is performed with adequate precision to be lossless to the image.”; also, para 47, “In the image output by exposure compensation 282, the region of interest in the scene remains as auto-exposed by the camera, while the rest of the image outside the region of interest is compensated to an exposure (referred to as scene exposure) as determined form the ambient light information.”).
Claim(s) 3-4, 7, 9, 12-13, 17, 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Molholm (U.S. Pub. No 20210099632) as modified by Ollila (U.S. Pub. No. 20220343529) and Hirvonen et al. (U.S. Pub. No 20250124667), further in view of Kuo et al. (U.S. Doc. No. 9317930).
Regarding claim 3, Molholm as modified by Ollila and Hirvonen discloses the system of claim 1, wherein at least one visual adjustment algorithm comprises an automatic white balance algorithm that adjusts a color balance of the frame.
However, in a similar field of endeavor, Kuo discloses the at least one visual adjustment algorithm comprises an automatic white balance algorithm that adjusts a color balance of the frame (col 202, “In one embodiment, the application of the GOC2 logic 3012 subsequent to color correction may provide for auto-white balance of the image data based on the corrected color values, and may also adjust sensor variations of the red-to-green and blue-to-green ratios.”; also, col 62, “When a white object is illuminated under a low color temperature, it may appear reddish in the captured image. Conversely, a white object that is illuminated under a high color temperature may appear bluish in the captured image. The goal of white balancing is, therefore, to adjust RGB values such that the image appears to the human eye as if it were taken under canonical light”; also, col 141, “The WBG logic 1036 applies white balance gains to all four components at each pixel.”; also, col 198, “The values for G[c] may be previously determined during statistics processing.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Molholm in view of Ollila, further in view of Hirvonen, in which a mixed reality processing pipeline receives a frame of passthrough video, identifies the part of that frame over which superimposed content fully obscures the real world scene, derives the camera control values from the remaining portion, and adjusts the frame by an exposure compensation applied in the display pipeline, with the features of Kuo's invention of an automatic white balance that applies per component gains to the pixel values of the image data so that the colour rendered by the frame is changed. The combination would have been obvious because the base combination adjusts only the exposure of the displayed passthrough frame and leaves its colour uncorrected, and Kuo supplies the missing process, applying white balance gains to each pixel of the image data and stating that the goal of white balancing is to adjust the values so that the image appears to the human eye as if it were taken under canonical light. A person of ordinary skill would have been led to add it because Kuo derives the gains it applies from the same statistics collection the base combination already performs, and because a frame whose exposure has been corrected from the portion of the scene that is not obscured will still carry the colour cast of the illuminant unless a white balance is applied to the frame itself. The results would have been predictable to one of ordinary skill, who would have had a reasonable expectation of success, because Kuo applies the gains in a stage of the same image signal processing pipeline that already carries the frame, and because Molholm's gaze-positioned region of interest, its camera auto-exposure, its exposure compensation, and its blending operation each continue to operate as Molholm describes.
Regarding claim 4, Molholm as modified by Ollila and Hirvonen discloses the system of claim 1, wherein at least one visual adjustment algorithm comprises an automatic focus algorithm that adjusts a focus of the frame.
However, in a similar field of endeavor, Kuo discloses the at least one visual adjustment algorithm comprises an automatic focus algorithm that adjusts a focus of the frame (col 63, “Further, auto-focus may refer to determining the optimal focal length of the lens in order to substantially optimize the focus of the image. In certain embodiments, floating windows of high frequency statistics may be collected and the focal length of the lens may be adjusted to bring an image into focus. As discussed further below, in one embodiment, auto-focus adjustments may use coarse and fine adjustments based upon one or more metrics, referred to as auto-focus scores (AF scores) to bring an image into focus.”; also, col 81, “Using the AF statistics, the ISP control logic 84 (FIG. 7) may be configured to adjust a focal length of the lens of an image device (e.g., 30) using a series of focal length adjustments based on coarse and fine auto-focus "scores" to bring an image into focus.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Molholm in view of Ollila, further in view of Hirvonen, in which a mixed reality processing pipeline receives a frame of passthrough video, identifies the part of that frame over which superimposed content fully obscures the real world scene, derives the camera control values from the remaining portion, and adjusts the frame by an exposure compensation applied in the display pipeline, with the features of Kuo's invention of an auto-focus algorithm that collects high frequency statistics over windows of the image data and adjusts the focal length of the lens through a series of coarse and fine adjustments driven by auto-focus scores. The combination would have been obvious because the base combination sets the exposure of the displayed passthrough frame from the portion of the field of view that represents the physical scene but leaves the sharpness of that frame uncontrolled, and Kuo supplies the missing process, collecting its focus statistics from the same captured image data on which the base combination already computes its exposure statistics. A person of ordinary skill would have been led to add it because Kuo groups auto-focus with auto-exposure as statistics that a single image signal processor collects from one captured stream, so an artisan who had already routed the portion of the frame that is not obscured to the auto-exposure statistics would have routed it to the auto-focus statistics as well. The results would have been predictable to one of ordinary skill, who would have had a reasonable expectation of success, because Kuo derives the auto-focus score from ordinary high frequency filtering of the same image data the base combination already carries, and because Molholm's gaze-positioned region of interest, its camera auto-exposure, its exposure compensation, and its blending operation each continue to operate as Molholm describes.
Regarding claim 7, Molholm as modified by Ollila and Hirvonen discloses the system of claim 1, wherein modified version of the frame of the passthrough video without the region that is occluded by the virtual object comprises applying a weight to the region based on a transparency of the virtual object.
However, in a similar field of endeavor, Kuo discloses creating the modified version of the frame of the passthrough video without the region that is occluded by the virtual object comprises applying a weight to the region (col 48, “The statistics core 146a may include "3A" statistics collection logic 482 to collect statistics relating to auto-exposure, auto-white balance, auto-focus, and similar operations; fixed pattern noise (FPN) statistics collection logic 484; histogram statistics collection logic 486; and/or local statistics collection logic 488. “; also, col 74, “Further, by knowing the luma/AE statistics for tile statistics and window locations, AE metering may be performed. For instance, depending on the image scene, it may be desirable to weigh AE statistics at the center window more heavily than those at the edges of the image, such as may be in the case of a portrait. “; also, col 72, “Note also that the Count may be incremented by the weights such that the Count can be used for computing the weighted average values from the sums.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Molholm in view of Ollila, further in view of Hirvonen, in which a mixed reality processing pipeline identifies the part of a passthrough frame over which superimposed content obscures the real world scene and derives its camera control values from the remaining portion, with the features of Kuo's invention of assigning a weight to image data so that the weighted value is accumulated into the statistics sums and the pixel count is incremented by that same weight. The combination would have been obvious because the base combination treats the obscured part of the frame as a binary matter, either present in the derivation or absent from it, and Kuo supplies the graded alternative the base combination lacks, weighting how strongly a given portion of the frame contributes to the statistic rather than only whether it contributes at all. A person of ordinary skill would have been led to that alternative because Kuo applies it to the same three statistics the base combination computes, stating that its statistics core collects statistics relating to auto-exposure, auto-white balance and auto-focus, and because Kuo expressly contemplates weighting one window of the frame more heavily than another according to what the scene requires. The results would have been predictable to one of ordinary skill, who would have had a reasonable expectation of success, because Kuo increments the pixel count by the weight so that a weighted region is correctly normalized in the resulting average, and because the base combination already identifies the part of the frame to which a weight would be applied.
Ollila discloses based on a transparency of the virtual object (paragraph 58, “In some implementations, the blend object to be superimposed is a transparent blend object or a translucent blend object. Upon superimposition of such a blend object over the image segment, the image segment would only be partially obscured and a portion of the real-world scene corresponding to the image segment would still be visible in the at least one image. In such a case, when the given image signal lies in the at least one region, the given image signal is to be lightly processed to adequately represent said portion of the real-world scene in the at least one image.”; also, para 72, “Coefficients in the given colour conversion matrix may depend on a transparent blending region between the blend object and the image segment of the at least one image. In such a case, these coefficients could be estimated using a linear interpolation, based on alpha layer (having transparency values of pixels) in the transparent blending region.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Molholm in view of Ollila, further in view of Hirvonen, further in view of Kuo, in which a weight is applied to the part of a passthrough frame that superimposed content obscures so that the part contributes to the camera control statistics in proportion to that weight, with the features of Ollila's invention of setting the treatment of that part according to whether the superimposed object is opaque or is transparent or translucent, and of deriving a numerical coefficient by linear interpolation from an alpha layer carrying the transparency values of pixels. The combination would have been obvious because Kuo supplies a weight but says nothing about what should set it, and Ollila answers that question for this precise situation, teaching that where the superimposed object is transparent or translucent the image segment is only partially obscured and a portion of the real-world scene corresponding to it is still visible and therefore still warrants representation, whereas where the object is opaque that portion is not visible at all. A person of ordinary skill would have recognized that a portion of the scene which remains partly visible should contribute partly, and that the transparency of the object over it is the measure of how much remains visible. The results would have been predictable to one of ordinary skill, who would have had a reasonable expectation of success, because Ollila already derives numerical processing coefficients from per pixel transparency values by linear interpolation, so the transparency value needed to set Kuo's weight is a quantity the base combination already computes.
Regarding claim 9, Molholm as modified by Ollila and Hirvonen discloses the method of claim 8, wherein modified version of the frame of the passthrough video without the region that is occluded by the virtual object comprises applying a weight to the region based on a transparency of the virtual object
However, in a similar field of endeavor, Kuo discloses creating the modified version of the frame of the passthrough video without the region that is occluded by the virtual object comprises applying a weight to the region (col 48, “The statistics core 146a may include "3A" statistics collection logic 482 to collect statistics relating to auto-exposure, auto-white balance, auto-focus, and similar operations; fixed pattern noise (FPN) statistics collection logic 484; histogram statistics collection logic 486; and/or local statistics collection logic 488. “; also, col 74, “Further, by knowing the luma/AE statistics for tile statistics and window locations, AE metering may be performed. For instance, depending on the image scene, it may be desirable to weigh AE statistics at the center window more heavily than those at the edges of the image, such as may be in the case of a portrait. “; also, col 72, “Note also that the Count may be incremented by the weights such that the Count can be used for computing the weighted average values from the sums.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Molholm in view of Ollila, further in view of Hirvonen, in which a mixed reality processing pipeline identifies the part of a passthrough frame over which superimposed content obscures the real world scene and derives its camera control values from the remaining portion, with the features of Kuo's invention of assigning a weight to image data so that the weighted value is accumulated into the statistics sums and the pixel count is incremented by that same weight. The combination would have been obvious because the base combination treats the obscured part of the frame as a binary matter, either present in the derivation or absent from it, and Kuo supplies the graded alternative the base combination lacks, weighting how strongly a given portion of the frame contributes to the statistic rather than only whether it contributes at all. A person of ordinary skill would have been led to that alternative because Kuo applies it to the same three statistics the base combination computes, stating that its statistics core collects statistics relating to auto-exposure, auto-white balance and auto-focus, and because Kuo expressly contemplates weighting one window of the frame more heavily than another according to what the scene requires. The results would have been predictable to one of ordinary skill, who would have had a reasonable expectation of success, because Kuo increments the pixel count by the weight so that a weighted region is correctly normalized in the resulting average, and because the base combination already identifies the part of the frame to which a weight would be applied.
Ollila discloses based on a transparency of the virtual object (paragraph 58, “In some implementations, the blend object to be superimposed is a transparent blend object or a translucent blend object. Upon superimposition of such a blend object over the image segment, the image segment would only be partially obscured and a portion of the real-world scene corresponding to the image segment would still be visible in the at least one image. In such a case, when the given image signal lies in the at least one region, the given image signal is to be lightly processed to adequately represent said portion of the real-world scene in the at least one image.”; also, para 72, “Coefficients in the given colour conversion matrix may depend on a transparent blending region between the blend object and the image segment of the at least one image. In such a case, these coefficients could be estimated using a linear interpolation, based on alpha layer (having transparency values of pixels) in the transparent blending region.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Molholm in view of Ollila, further in view of Hirvonen, further in view of Kuo, in which a weight is applied to the part of a passthrough frame that superimposed content obscures so that the part contributes to the camera control statistics in proportion to that weight, with the features of Ollila's invention of setting the treatment of that part according to whether the superimposed object is opaque or is transparent or translucent, and of deriving a numerical coefficient by linear interpolation from an alpha layer carrying the transparency values of pixels. The combination would have been obvious because Kuo supplies a weight but says nothing about what should set it, and Ollila answers that question for this precise situation, teaching that where the superimposed object is transparent or translucent the image segment is only partially obscured and a portion of the real-world scene corresponding to it is still visible and therefore still warrants representation, whereas where the object is opaque that portion is not visible at all. A person of ordinary skill would have recognized that a portion of the scene which remains partly visible should contribute partly, and that the transparency of the object over it is the measure of how much remains visible. The results would have been predictable to one of ordinary skill, who would have had a reasonable expectation of success, because Ollila already derives numerical processing coefficients from per pixel transparency values by linear interpolation, so the transparency value needed to set Kuo's weight is a quantity the base combination already computes.
Regarding claim 12, Molholm as modified by Ollila and Hirvonen discloses the method of claim 8, wherein at least one visual adjustment algorithm comprises an automatic white balance algorithm that adjusts a color balance of the frame.
However, in a similar field of endeavor, Kuo discloses the at least one visual adjustment algorithm comprises an automatic white balance algorithm that adjusts a color balance of the frame (col 202, “In one embodiment, the application of the GOC2 logic 3012 subsequent to color correction may provide for auto-white balance of the image data based on the corrected color values, and may also adjust sensor variations of the red-to-green and blue-to-green ratios.”; also, col 62, “When a white object is illuminated under a low color temperature, it may appear reddish in the captured image. Conversely, a white object that is illuminated under a high color temperature may appear bluish in the captured image. The goal of white balancing is, therefore, to adjust RGB values such that the image appears to the human eye as if it were taken under canonical light”; also, col 141, “The WBG logic 1036 applies white balance gains to all four components at each pixel.”; also, col 198, “The values for G[c] may be previously determined during statistics processing.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Molholm in view of Ollila, further in view of Hirvonen, in which a mixed reality processing pipeline receives a frame of passthrough video, identifies the part of that frame over which superimposed content fully obscures the real world scene, derives the camera control values from the remaining portion, and adjusts the frame by an exposure compensation applied in the display pipeline, with the features of Kuo's invention of an automatic white balance that applies per component gains to the pixel values of the image data so that the colour rendered by the frame is changed. The combination would have been obvious because the base combination adjusts only the exposure of the displayed passthrough frame and leaves its colour uncorrected, and Kuo supplies the missing process, applying white balance gains to each pixel of the image data and stating that the goal of white balancing is to adjust the values so that the image appears to the human eye as if it were taken under canonical light. A person of ordinary skill would have been led to add it because Kuo derives the gains it applies from the same statistics collection the base combination already performs, and because a frame whose exposure has been corrected from the portion of the scene that is not obscured will still carry the colour cast of the illuminant unless a white balance is applied to the frame itself. The results would have been predictable to one of ordinary skill, who would have had a reasonable expectation of success, because Kuo applies the gains in a stage of the same image signal processing pipeline that already carries the frame, and because Molholm's gaze-positioned region of interest, its camera auto-exposure, its exposure compensation, and its blending operation each continue to operate as Molholm describes.
Regarding claim 13, Molholm as modified by Ollila and Hirvonen discloses the method of claim 8, wherein at least one visual adjustment algorithm comprises an automatic focus algorithm that adjusts a focus of the frame.
However, in a similar field of endeavor, Kuo discloses the at least one visual adjustment algorithm comprises an automatic focus algorithm that adjusts a focus of the frame (col 63, “Further, auto-focus may refer to determining the optimal focal length of the lens in order to substantially optimize the focus of the image. In certain embodiments, floating windows of high frequency statistics may be collected and the focal length of the lens may be adjusted to bring an image into focus. As discussed further below, in one embodiment, auto-focus adjustments may use coarse and fine adjustments based upon one or more metrics, referred to as auto-focus scores (AF scores) to bring an image into focus.”; also, col 81, “Using the AF statistics, the ISP control logic 84 (FIG. 7) may be configured to adjust a focal length of the lens of an image device (e.g., 30) using a series of focal length adjustments based on coarse and fine auto-focus "scores" to bring an image into focus.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Molholm in view of Ollila, further in view of Hirvonen, in which a mixed reality processing pipeline receives a frame of passthrough video, identifies the part of that frame over which superimposed content fully obscures the real world scene, derives the camera control values from the remaining portion, and adjusts the frame by an exposure compensation applied in the display pipeline, with the features of Kuo's invention of an auto-focus algorithm that collects high frequency statistics over windows of the image data and adjusts the focal length of the lens through a series of coarse and fine adjustments driven by auto-focus scores. The combination would have been obvious because the base combination sets the exposure of the displayed passthrough frame from the portion of the field of view that represents the physical scene but leaves the sharpness of that frame uncontrolled, and Kuo supplies the missing process, collecting its focus statistics from the same captured image data on which the base combination already computes its exposure statistics. A person of ordinary skill would have been led to add it because Kuo groups auto-focus with auto-exposure as statistics that a single image signal processor collects from one captured stream, so an artisan who had already routed the portion of the frame that is not obscured to the auto-exposure statistics would have routed it to the auto-focus statistics as well. The results would have been predictable to one of ordinary skill, who would have had a reasonable expectation of success, because Kuo derives the auto-focus score from ordinary high frequency filtering of the same image data the base combination already carries, and because Molholm's gaze-positioned region of interest, its camera auto-exposure, its exposure compensation, and its blending operation each continue to operate as Molholm describes.
Regarding claim 17, Molholm as modified by Ollila and Hirvonen discloses the non-transitory computer-readable medium of claim 16, wherein modified version of the frame of the passthrough video without the region that is occluded by the virtual object comprises applying a weight to the region based on a transparency of the virtual object.
However, in a similar field of endeavor, Kuo discloses creating the modified version of the frame of the passthrough video without the region that is occluded by the virtual object comprises applying a weight to the region (col 48, “The statistics core 146a may include "3A" statistics collection logic 482 to collect statistics relating to auto-exposure, auto-white balance, auto-focus, and similar operations; fixed pattern noise (FPN) statistics collection logic 484; histogram statistics collection logic 486; and/or local statistics collection logic 488. “; also, col 74, “Further, by knowing the luma/AE statistics for tile statistics and window locations, AE metering may be performed. For instance, depending on the image scene, it may be desirable to weigh AE statistics at the center window more heavily than those at the edges of the image, such as may be in the case of a portrait. “; also, col 72, “Note also that the Count may be incremented by the weights such that the Count can be used for computing the weighted average values from the sums.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Molholm in view of Ollila, further in view of Hirvonen, in which a mixed reality processing pipeline identifies the part of a passthrough frame over which superimposed content obscures the real world scene and derives its camera control values from the remaining portion, with the features of Kuo's invention of assigning a weight to image data so that the weighted value is accumulated into the statistics sums and the pixel count is incremented by that same weight. The combination would have been obvious because the base combination treats the obscured part of the frame as a binary matter, either present in the derivation or absent from it, and Kuo supplies the graded alternative the base combination lacks, weighting how strongly a given portion of the frame contributes to the statistic rather than only whether it contributes at all. A person of ordinary skill would have been led to that alternative because Kuo applies it to the same three statistics the base combination computes, stating that its statistics core collects statistics relating to auto-exposure, auto-white balance and auto-focus, and because Kuo expressly contemplates weighting one window of the frame more heavily than another according to what the scene requires. The results would have been predictable to one of ordinary skill, who would have had a reasonable expectation of success, because Kuo increments the pixel count by the weight so that a weighted region is correctly normalized in the resulting average, and because the base combination already identifies the part of the frame to which a weight would be applied.
Ollila discloses based on a transparency of the virtual object (paragraph 58, “In some implementations, the blend object to be superimposed is a transparent blend object or a translucent blend object. Upon superimposition of such a blend object over the image segment, the image segment would only be partially obscured and a portion of the real-world scene corresponding to the image segment would still be visible in the at least one image. In such a case, when the given image signal lies in the at least one region, the given image signal is to be lightly processed to adequately represent said portion of the real-world scene in the at least one image.”; also, para 72, “Coefficients in the given colour conversion matrix may depend on a transparent blending region between the blend object and the image segment of the at least one image. In such a case, these coefficients could be estimated using a linear interpolation, based on alpha layer (having transparency values of pixels) in the transparent blending region.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Molholm in view of Ollila, further in view of Hirvonen, further in view of Kuo, in which a weight is applied to the part of a passthrough frame that superimposed content obscures so that the part contributes to the camera control statistics in proportion to that weight, with the features of Ollila's invention of setting the treatment of that part according to whether the superimposed object is opaque or is transparent or translucent, and of deriving a numerical coefficient by linear interpolation from an alpha layer carrying the transparency values of pixels. The combination would have been obvious because Kuo supplies a weight but says nothing about what should set it, and Ollila answers that question for this precise situation, teaching that where the superimposed object is transparent or translucent the image segment is only partially obscured and a portion of the real-world scene corresponding to it is still visible and therefore still warrants representation, whereas where the object is opaque that portion is not visible at all. A person of ordinary skill would have recognized that a portion of the scene which remains partly visible should contribute partly, and that the transparency of the object over it is the measure of how much remains visible. The results would have been predictable to one of ordinary skill, who would have had a reasonable expectation of success, because Ollila already derives numerical processing coefficients from per pixel transparency values by linear interpolation, so the transparency value needed to set Kuo's weight is a quantity the base combination already computes.
Regarding claim 20, Molholm as modified by Ollila and Hirvonen discloses the non-transitory computer-readable medium of claim 16, wherein at least one visual adjustment algorithm comprises an automatic white balance algorithm that adjusts a color balance of the frame.
However, in a similar field of endeavor, Kuo discloses the at least one visual adjustment algorithm comprises an automatic white balance algorithm that adjusts a color balance of the frame the at least one visual adjustment algorithm comprises an automatic white balance algorithm that adjusts a color balance of the frame (col 202, “In one embodiment, the application of the GOC2 logic 3012 subsequent to color correction may provide for auto-white balance of the image data based on the corrected color values, and may also adjust sensor variations of the red-to-green and blue-to-green ratios.”; also, col 62, “When a white object is illuminated under a low color temperature, it may appear reddish in the captured image. Conversely, a white object that is illuminated under a high color temperature may appear bluish in the captured image. The goal of white balancing is, therefore, to adjust RGB values such that the image appears to the human eye as if it were taken under canonical light”; also, col 141, “The WBG logic 1036 applies white balance gains to all four components at each pixel.”; also, col 198, “The values for G[c] may be previously determined during statistics processing.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have further modified Molholm in view of Ollila, further in view of Hirvonen, in which a mixed reality processing pipeline receives a frame of passthrough video, identifies the part of that frame over which superimposed content fully obscures the real world scene, derives the camera control values from the remaining portion, and adjusts the frame by an exposure compensation applied in the display pipeline, with the features of Kuo's invention of an automatic white balance that applies per component gains to the pixel values of the image data so that the colour rendered by the frame is changed. The combination would have been obvious because the base combination adjusts only the exposure of the displayed passthrough frame and leaves its colour uncorrected, and Kuo supplies the missing process, applying white balance gains to each pixel of the image data and stating that the goal of white balancing is to adjust the values so that the image appears to the human eye as if it were taken under canonical light. A person of ordinary skill would have been led to add it because Kuo derives the gains it applies from the same statistics collection the base combination already performs, and because a frame whose exposure has been corrected from the portion of the scene that is not obscured will still carry the colour cast of the illuminant unless a white balance is applied to the frame itself. The results would have been predictable to one of ordinary skill, who would have had a reasonable expectation of success, because Kuo applies the gains in a stage of the same image signal processing pipeline that already carries the frame, and because Molholm's gaze-positioned region of interest, its camera auto-exposure, its exposure compensation, and its blending operation each continue to operate as Molholm describes.
Response to Arguments
Applicant's arguments filed 07/01/2026 have been fully considered.
Applicant argues at pages 12 to 14 of the Remarks that Molholm discloses gaze-driven region of interest statistics and exposure compensation but is silent on using a frame modified to exclude occluded regions to calculate statistics on the modified frame. This argument is persuasive. Molholm gathers image statistics from a spot metering region positioned on the full image from the camera, and paragraph [0022] states that "The position of the ROI (Region of Interest) on the full image from the camera is based on the user's gaze direction as determined by the eye tracking system." Molholm further computes the region of interest statistics before the frame exists, paragraph [0046] stating that "The ROI statistics are provided to sensor gain 252 so that an image is captured by camera 250 that is auto-exposed for a region of interest in a scene determined from the point of gaze based on a metered result through a combination of integration time and gain in order to acquire a properly exposed image (with the least amount of noise) within the ROI." No version of a received frame is supplied to an analysis algorithm in Molholm. The rejection of the previous action to the extent it relied upon Molholm for the calculation of a statistic based on the modified version of the frame is withdrawn.
Applicant's characterization of Molholm as directed to an area of a frame the user is currently focused on is persuasive in a second respect. The previous action relied upon Molholm at paragraphs [0018] and [0023] for the automatic exposure control dependent claims, and each of those passages is limited to the region of interest rather than to the frame, paragraph [0018] describing "selective auto-exposure for a region of interest in a scene based on gaze" and paragraph [0023] describing an image auto-exposed "within the ROI." Neither passage shows an exposure level of the frame being adjusted. Those mappings are withdrawn and the automatic exposure control claims are now rejected on Molholm at paragraph [0047], which applies exposure compensation to the image itself and compensates the portion of the image outside the region of interest. The automatic white balance and automatic focus dependent claims are likewise no longer rejected on the passages relied upon in the previous action.
Upon further consideration, new grounds of rejection under 35 U.S.C. 103 over Molholm in view of Ollila and Hirvonen, and Molholm in view of Ollila and Hirvonen, further in view of Kuo, are made as set forth.
Applicant argues at pages 13 to 14 of the Remarks that Lei is directed to identifying a general region of interest and tracking the user's real world physical hands relative to that region of interest, and that this is distinct from detecting an area of a frame occluded by a virtual object. This argument is persuasive. Lei detects the user's physical hands by means of the depth sensor 320 or the front facing sensor 1525 and compares their location to a region of interest or to the known location of a hologram, and the response of Lei is to trigger or to suppress the auto focus operation rather than to modify a frame. Lei is withdrawn from the rejection and is not applied in this action.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jai Li whose telephone number is (571)272-1170. The examiner can normally be reached Mon-Thu between 06:00-16:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached at (571)272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JAI W LI/Junior Examiner, Art Unit 2613
/XIAO M WU/Supervisory Patent Examiner, Art Unit 2613