DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 10/07/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
CLAIM INTERPRETATION
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
§ 112(f) interpretation despite the absence of “means.”
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: the “acquirer” and “generator” for claims 1-6.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1,2,3 is/are rejected under 35 U.S.C. 103 as being unpatentable over Kim(US20190347501A1) in view of Raffle(US-20190347501-A1).
As per claim 1, Kim teaches “an information processing apparatus comprising:” ( Kim further teaches at para. [0021], “The exemplary electronic device 100 that may be any type of device having multiple cameras. The electronic device 100 may comprise a processor 102 coupled to a plurality of cameras 104 and 105. While cameras 104 and 105 are shown, it should be understood that the cameras comprise image recording devices, such as image sensors, and that the cameras may be independent of each other or may share circuitry. The cameras may a part of an HMD, where one camera may be used for providing a view of a scene and the other may be used for performing eye tracking (i.e. the movement and viewing direction) of the eyes of a user of the HMD, as will be described in more detail below. The mobile device 100 could be any type of device adapted to transmit and receive information, such as a smart phone, tablet or other electronic device receiving or providing information, such as a wearable device. The processor 102 is an integrated electronic circuit such as, for example, an ARM processor, an X86 processor, a MIPS processor, a graphics processing unit (GPU), a general purpose GPU, or any other processor configured to execute instructions stored in a memory. ”); “an acquirer configured to acquire image information indicating an image captured by the image capture device” (Kim teaches at para. [0025], “As shown in FIG. 2, an image recording circuit 202, which may be one of the cameras of FIG. 1 for example, is used to record an image that is provided to a crop and resize block 204 to crop the image to a smaller area of the original image and to have relevant areas of the original image .” This disclosure teaches acquisition, by the processor/electronic device, of image information indicating an image captured by the image capture device.); “an image capture device is mounted at the head of the user” (Kim teaches at para. [0028], “The head-mounted electronic device 300 of FIG. 3 comprises a head mounting element 302 enabling securing a control unit 304 to a user's head. ” Kim further teaches at para. [0029], “A camera 310, which may be a part of the portable electronic device 306, allows the head-mounted electronic device to function as a virtual reality (VR) device or an augmented reality (AR) device using the camera to pass-through images of the surrounding environment. A second camera or other eye tracking sensor may be employed inside the HMD to perform eye tracking. ” Kim also teaches at para. [0021] that “The cameras may a part of an HMD, where one camera may be used for providing a view of a scene and the other may be used for performing eye tracking (i.e. the movement and viewing direction) of the eyes of a user of the HMD, as will be described in more detail below “ These disclosures teach that the image capture device is part of the HMD worn on the user's head.); “a generator configured to control a position in the captured image, at which the captured image is partially cropped” (Kim teaches at para. [0031], “The eye tracking block 404 allows for identifying a location where a user of the HMD is looking, which may be represented by X and Y coordinates for example. The X and Y coordinates may be used to position a bounding box used for cropping the image. ” Kim further teaches at para. [0033], “The crop and resize block 204 receives recorded images, as well as eye tracking information from the eye tracking block 404 and depth information from the depth camera block 406, to determine a how to crop an image. ” These disclosures teach controlling the image-space position of a cropping region or bounding box within the captured image.); “generate, by controlling the position in the captured image, a partial image cropped from the captured image” Kim teaches at para. [0025] “As shown in FIG. 2, an image recording circuit 202, which may be one of the cameras of FIG. 1 for example, is used to record an image that is provided to a crop and resize block 204 to crop the image to a smaller area of the original image and to have relevant areas of the original image. ” Kim further teaches at para. [0041], “A region of interest of an image based upon the eye tracking is determined at a block 1104. A bounding box is generated based upon the region of interest at a block 1106. The size of the bounding box may also be generated based upon depth information associated with objects in the image. An image is then cropped based upon the bounding box to generate a cropped image at a block 1108.” These disclosures teach generating a partial image from the captured image by cropping the image according to the controlled bounding-box position.); and “an image processor configured to perform image processing on the partial image” (Kim teaches at para. [0027], “The cropped image is then provided to a deep learning block 206. The deep learning block 206 performs deep learning that, unlike task-specific processing, makes decisions or provides outputs based upon the relationship between various detected stimuli or conditions.” Kim further teaches that “Deep learning can be used in object detection to not only identify particular objects based upon other objects in a scene, but to also determine the relevance of a particular object in a scene or the relationship between different objects in the scene.” Kim further teaches at para. [0033], “The cropped image is provided to the deep learning block 206 to generate object detection information,” and at para. [0042], “detecting an object in the cropped image may comprise performing deep learning.” These disclosures teach performing image processing, including deep-learning-based object detection, on the partial cropped image rather than on the full captured image.).
However, Kim does not expressly disclose “movement information about movement of a user” in the conservative sense of movement of the user's head or corresponding movement of the head-mounted apparatus, nor does Kim expressly disclose that the position at which the image is cropped is controlled “in accordance with the movement information.” Raffle supplies these limitations. Raffle teaches “movement information about movement of a user” (Raffle teaches at para. [0056], “For example, when HMD 102 is worn, HMD 102 may use one or more gyroscopes and/or one or more accelerometers to detect head movement. The HMD 102 may then interpret certain head-movements as being user input, such as nodding, or looking up, down, left, or right. An HMD 102 could also pan or scroll through graphics in a display according to movement. Other types of actions may also be mapped to head movement.” This disclosure teaches acquiring movement information corresponding specifically to movement of the user's head while wearing the HMD.); and Raffle teaches controlling an image position “in accordance with the movement information” (Raffle teaches at para. [0036], “X-axis and Y-axis movements can be performed using UI events that track head movements of the wearer. For example, ‘panning’ or moving up in the display of the object can be directed by an upward head movement by the wearer. Similarly, panning down, left, and right within the display of the object can be performed by using suitable head movements.” Raffle further teaches at para. [0039], “For example, an X-axis movement of X° can act as an instruction to the ZAOD to move a display by a number of pixels NPx along the X-axis based on the X-axis movement such that NPx is proportional to the Z axis component ,” and similarly, “a Y-axis movement of Y° can act as an instruction to the ZAOD to move a display by a number of pixels NPy along the Y-axis based on the Y-axis movement such that NPy is proportional to Z.” Raffle further relates the resulting image position to a crop region, teaching at para. [0037], “Once navigation is complete, the ‘cropped’ object or object as displayed can be saved for future use, such as sharing with others, use with other applications,” and “The cropped object can be specified using a ‘cropping window’ of the original object.” ). It would have been obvious to a person of ordinary skill in the art before the effective filing date to modify Kim's HMD image-cropping system so that the image-space position of Kim's ROI/bounding box is additionally adjusted in accordance with head-movement information using Raffle's head-movement-to-pixel-displacement technique. head-mounted camera changed orientation or position, Kim's ROI could be correspondingly repositioned in image space to remain associated with the intended object or scene region. Such a modification would have been a predictable application of known motion-sensing and image-coordinate adjustment techniques to Kim's already-disclosed crop-and-resize process, with the predictable benefit of maintaining Kim's ROI over the desired target despite movement of the HMD wearer, thereby rendering obvious the claimed control of the crop position in accordance with movement information.
As per claim 2, Kim teaches “the generator is configured to generate the partial image by cropping a region showing an object previously designated” (Kim teaches at para. [0041], “A region of interest of an image based upon the eye tracking is determined at a block 1104. A bounding box is generated based upon the region of interest at a block 1106. ... An image is then cropped based upon the bounding box to generate a cropped image at a block 1108.” This teaches identifying a region of interest corresponding to an object, generating a bounding box for that region, and cropping the image according to the previously determined region.). However, Kim does not expressly disclose “the acquirer is configured to acquire, as the movement information, information about movement of the head of the user” or generating the cropped region “in accordance with the information about the movement of the head.” Raffle supplies these limitations. Raffle teaches “the acquirer is configured to acquire, as the movement information, information about movement of the head of the user” (Raffle teaches at para. [0056], “For example, when HMD 102 is worn, HMD 102 may use one or more gyroscopes and/or one or more accelerometers to detect head movement. The HMD 102 may then interpret certain head-movements as being user input, such as nodding, or looking up, down, left, or right. An HMD 102 could also pan or scroll through graphics in a display according to movement. Other types of actions may also be mapped to head movement. ” This expressly teaches acquiring head-movement information.); and Raffle teaches generating the cropped region “in accordance with the information about the movement of the head” (Raffle teaches at para. [0038], “an X-axis movement of X° can act as an instruction to the ZAOD to move a display by a number of pixels NPx along the X-axis based on the X-axis movement,” and similarly “a Y-axis movement of Y° can act as an instruction to the ZAOD to move a display by a number of pixels NPy along the Y-axis based on the Y-axis movement.” Raffle further teaches at para. [0037], “The cropped object can be specified using a ‘cropping window’ of the original object.” These disclosures teach using detected head movement to determine image-space displacement and thereby control the position of the cropping window.) The rationale of claim 1 is incorporated herein
As per claim 3, Kim does not expressly disclose “the acquirer is configured to acquire the information about the movement of the head, from an inertia measurement device or a geomagnetic sensor mounted at the head of the user”.
Raffle teaches “the acquirer is configured to acquire the information about the movement of the head, from an inertia measurement device or a geomagnetic sensor mounted at the head of the user” (Raffle teaches at para. [0052] that the HMD may include “one or more gyroscopes, one or more accelerometers, one or more magnetometers,” and at para. [0056] that “when HMD 102 is worn, HMD 102 may use one or more gyroscopes and/or one or more accelerometers to detect head movement.” These disclosures teach acquiring head-movement information from inertial sensors mounted on the HMD, and the disclosed magnetometer corresponds to the claimed geomagnetic sensor alternative.) The rationale of claim 1 is incorporated herein
Claim(s) 4,5,6 are rejected under 35 U.S.C. 103 as being unpatentable over Kim as modified by Raffle as applied to claim 1 above, and further in view of Tall(US-20190318708-A1).
As per claim 4, Kim teaches “the acquirer is configured to acquire, as the movement information, eye information about movement of eyes of the user” (Kim teaches at para. [0021] that one camera of the HMD may be used for “performing eye tracking (i.e. the movement and viewing direction) of the eyes of a user of the HMD.” Kim further teaches at para. [0031] that “The eye tracking block 404 allows for identifying a location where a user of the HMD is looking, which may be represented by X and Y coordinates for example.” These disclosures teach acquiring eye-movement and gaze-direction information of the HMD wearer.). However, Kim does not expressly disclose “the generator is configured to generate the partial image by cropping from the captured image a region out of a region in which predetermined information is recognizable by the user, based on the eye information.”
Tall teaches“the generator is configured to generate the partial image by cropping from the captured image a region out of a region in which predetermined information is recognizable by the user, based on the eye information” (Tall teaches at para.[0022] “The display system identifies the first region and the second region based on eye tracking information received from an eye tracking unit. The display system uses the eye tracking information to determine the point on the screen at which the user is looking (hereinafter referred to as the point of regard). The display system can then determine the boundaries of the fovea region, the parafovea region, and the outside region based on the point of regard …. The outside region is the portion of the screen beyond the outside radius of the parafovea region.” Tall explains at para.[0023] the perceptual basis for selecting that outside region, stating that “the region of the screen around the point of regard (i.e., the fovea region) will appear to have higher image quality, and this is the region to which the eye is most sensitive to image quality, … in the outside region, lower parameters can be applied without noticeable image quality degradation.” Tall further teaches at para.[0073] that these eye-defined screen regions to image regions and cropping, teaching that “The system 110 renders and/or encodes an image by applying 406 different sets of parameters to different image regions within the image. In embodiments where the image is to be displayed on the entire screen, the image regions are coextensive with the screen regions. In embodiments where the image is to be displayed on a portion of the screen, the system 110 determines the boundaries of the image regions by cropping the screen regions to the portion of the screen on which the image is to be displayed. The display system 110 transmits 408 the image to the display device 105 to be displayed on the screen. ” These disclosures teach using eye-tracking information to determine a foveal/parafoveal region corresponding to the portion of the image where the user's visual sensitivity and recognition ability are greatest, determining an outside region beyond that visually sensitive region, and defining corresponding image-region boundaries by cropping. Thus, Tall teaches selecting, based on eye information, a region outside the region in which visual information is most readily recognizable by the user.) It would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to to modify the combined system of Kim and Raffle with Tall’s gaze-based partitioning so that Kim’s crop region is selected from an outside/peripheral region identified from the user’s eye information, while Raffle’s head-movement displacement technique maintains that region as the HMD wearer moves. This combination would predictably focus image-processing resources on portions of the captured image that are less visually recognizable to the user while preserving the selected ROI during head movement, thereby improving processing efficiency.
As per claim 5, the combination of Kim, Raffle and Tall disclose all the elements of claim 4 as discussed above. Tall also discloses “the generator is configured to identify a position of a viewpoint of the user, based on the eye information” (Tall teaches at para. [0064] that “The screen region module 162 determines a gaze vector 306 representing the direction in which the eye is looking. … After determining the gaze vector 306, the module 162 determines the point of regard 314 by computing an intersection between the gaze vector 306 and the screen 304.” Tall also explains that the gaze vector may be derived from eye-tracking information, including the angular orientation or foveal axis of the user's eye. These disclosures teach determining the user's point of regard, i.e., viewpoint, based on eye information.) Tall further teaches “the generator is configured to crop as the partial image a portion separated from the viewpoint by a predetermined distance or more” (Tall teaches at para. [0066] that “the screen region module 162 determines the fovea region 308 by drawing a circle of a predetermined radius centered on the point of regard 314. Similarly, the screen region module 162 determines the parafovea region 310 by drawing a second circle of a second predetermined radius centered on the point of regard 314 and subtracting the foveal region 308 to yield an annulus (a ring shape). The screen region module 162 may additionally determine the outside region 312 as the region of the screen not covered by the fovea region 308 and the parafovea region 310.” Tall further teaches at para. [0067] that “the predetermined radii may be selected to correspond to a particular number of degrees of visual angle,” and gives exemplary radii of “2.5 degrees of visual angle” for the fovea region and “5.0 degrees of visual angle” for the parafovea region. Tall further teaches at para. [0074] that “the system 110 determines the boundaries of the image regions by cropping the screen regions to the portion of the screen on which the image is to be displayed.” These disclosures teach defining, relative to the user's point of regard, an outside region that begins beyond a predetermined radius and defining corresponding image-region boundaries by cropping; thus, Tall teaches selecting a portion separated from the user's viewpoint by at least a predetermined distance.). The rationale of claim 4 is incorporated herein.
As per claim 6, the combination of Kim, Raffle and Tall disclose all the elements of claim 4 as discussed above. The combination also discloses “the partial image includes a first partial image and a second partial image” (Tall teaches at para. [0058] that “The module 164 renders images by dividing the images into two or more image regions and applying a corresponding set of rendering parameters to each region.” This teaches dividing an image into at least first and second image regions, which correspond to first and second partial images of the overall image;) “a degree of gazing at the first partial image by the user is different from a degree of gazing at the second partial image by the user” (Tall teaches at para. [0065] that “The region of the screen depicted as 302 is the fovea region where the eye would be most sensitive to differences in resolution. The region depicted as 304 is the parafovea region, closest to 302, where the eye is less sensitive to differences in resolution. The area outside regions 302 and 304 is the outside image region 306, where the eye is least sensitive to difference in resolution.” This teaches image regions associated with different degrees of visual attention or gaze, with the foveal region receiving the greatest degree of gaze and the parafoveal/outside regions receiving progressively lesser degrees of gaze;) “the generator is configured to identify a position of a viewpoint of the user, based on the eye information” (Tall teaches at para. [0064] that “The screen region module 162 determines a gaze vector 306 representing the direction in which the eye is looking … After determining the gaze vector 306, the module 162 determines the point of regard 314 by computing an intersection between the gaze vector 306 and the screen 304.” This teaches determining the user's viewpoint, i.e., point of regard, from eye-tracking information;)Tall teaches “the generator is configured to crop the first partial image and the second partial image, based on a distance from the position of the viewpoint” (Tall teaches at para. [0066] that “the screen region module 162 determines the fovea region 308 by drawing a circle of a predetermined radius centered on the point of regard 314. Similarly, the screen region module 162 determines the parafovea region 310 by drawing a second circle of a second predetermined radius centered on the point of regard 314 and subtracting the foveal region 308 to yield an annulus (a ring shape). The screen region module 162 may additionally determine the outside region 312 as the region of the screen not covered by the fovea region 308 and the parafovea region 310.” Tall further teaches at para. [0074] that “the system 110 determines the boundaries of the image regions by cropping the screen regions to the portion of the screen on which the image is to be displayed.” These disclosures teach defining multiple image regions according to their respective distances from the user's point of regard and cropping those regions to establish corresponding first and second partial-image regions;) “image processing to be performed on the first partial image by the image processor differs from that on the second partial image by the image processor” (Tall teaches at para. [0076] that “The display system 110 renders 502 a first image region based on a first set of rendering parameters. The display system 110 also renders 504 a second image region based on a second set of rendering parameters.” This teaches performing different image processing on respective first and second image regions by applying different rendering parameters to each region.) The rationale of claim 4 is incorporated herein.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CHRIS ALEJANDRO PUNTIER whose telephone number is (703)756-1893. The examiner can normally be reached M-F 7:30-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Hajnik can be reached at 571-272-7642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CHRIS ALEJANDRO PUNTIER/ Examiner, Art Unit 2616
/DANIEL F HAJNIK/ Supervisory Patent Examiner, Art Unit 2616