DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Election/Restrictions
Applicant’s election of Species II in the reply filed on 5/6/2026 is acknowledged. Because applicant did not distinctly and specifically point out the supposed errors in the restriction requirement, the election has been treated as an election without traverse (MPEP § 818.01(a)).
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 31-35, 40, 42, 45, 48, 52 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Hafstad et al. (US 2023/0269468 A1) hereinafter referenced as Hafstad.
Regarding claim 31, Hafstad discloses
A method for automatically selecting video output data (Stream selector 204, 305 selects a stream to be output as the updated main focus stream; [0049]; fig. 1), the video output data comprising at least a sequence of layouts, wherein a layout is a combination or a composition of one or more visual sources in one video output, the visual source being at least provided by a camera (100, 101; fig. 1) shot of a scene, of a plurality of different camera shots of the scene provided by a plurality of cameras in different viewpoints, the scene comprising at least one actor and an object of interest ([0041]-[0050]; A virtual director applies a rule set to decide which shot should be shown at any moment. This results in different camera angles and zoom levels depending on the various objects in the scene.), the method comprising the steps of
configuring a plurality of layouts (Layouts are configured by the stream selectors 205, 305; fig. 2),
determining a plurality of interaction states (The vision pipeline units 203, 302 processes the incoming video to detect postures, positions, orientations, gestures, etc. of objects; [0042]-[0043]), and associating to each interaction state a sequence of layouts comprising at least one camera shot for each interaction state (The virtual directors units 203, 303 causes the stream selectors 205, 305 to select a stream to be output based on the output from the vision pipeline units according to predetermined rules; [0048], [0049], wherein an interaction state is a situation in which an interaction between at least one actor and its environment remains unchanged and depends at least on a looking direction of the at least one actor ([0043]; Interaction states includes where objects are, the extent of their visibility in the view (This is includes the direction of the objects gaze and the changes in the objects gaze; [0054]-[0055]), if they are speaking or not, their facial expressions, their body positions and their head poses),
capturing the scene with the plurality of cameras to generate video data of the scene (Overview streams; fig. 1),
processing the video data of the scene to detect a current interaction state based on the looking direction of the at least one actor ([0042]-[0047], [0053]-[0055]; The overview stream is processed to determine a user’s gaze direction. The virtual director uses rules for stream selection based on the user’s gaze direction.), and
selecting the video output data (Updated main focus stream; fig. 1) to show the sequence of layouts corresponding to the current interaction state ([0048]-[0049]).
Regarding claim 32, Hafstad discloses everything claimed as applied above (see claim 1), in addition, Hafstad discloses, wherein the method further comprises the step of:
configuring a plurality of interaction state transitions (Predetermined rule sets for transitions are configured by the virtual director; [0053]-[0060]) between each interaction state, wherein each interaction state transition comprises a sequence of layouts comprising at least one camera shot,
processing the acquired video data of the scene to detect an interaction state transition (This is performed by the vision pipeline units 202, 302; [0042]-[0047]; fig. 1),
selecting the video output data to show the sequence of layouts corresponding to the interaction state transition (The virtual director units 203, 303 select the streams based on the findings of the vision pipeline units 202, 302; [0048]-0049]).
Regarding claim 33, Hafstad discloses everything claimed as applied above (see claim 31), in addition, Hafstad discloses, wherein the method further comprises the step of:
processing the acquired video data of the scene to detect a subsequent interaction state of the plurality of interaction states (This is performed by the vision pipeline units 202, 302; [0042]-[0047]; fig. 1),
selecting the video output data to show the sequence of layouts corresponding to the subsequent interaction state (The virtual director units 203, 303 select the streams based on the findings of the vision pipeline units 202, 302; [0048]-0049]).
Regarding claim 34, Hafstad discloses everything claimed as applied above (see claim 31), in addition, Hafstad discloses, wherein the sequence of layouts associated to an interaction state comprises a layout of different camera shots and/or video streams, and/or video sources ([0041]; Different streams are selected having different camera angles, and zoom levels. These streams are associated with the interaction states of the objects according to predetermined rules; [0048]-[0049], [0053]-[0060]).
Regarding claim 35, Hafstad discloses everything claimed as applied above (see claim 31), in addition, Hafstad discloses, wherein the video data of the scene is provided with audio data (Microphones are provided for associated audio; [0047]), and the method further comprises the step of performing speech detection of the audio data ([0043], [0047]; Objects are determined if they are speaking or not by the vision pipeline unit 202, 303.).
Regarding claim 40, Hafstad discloses everything claimed as applied above (see claim 31), in addition, Hafstad discloses, wherein the step of processing the acquired video data of the scene to detect a current interaction state comprises the step of evaluating the looking direction of the at least one actor ([0042]-[0047], [0053]-[0055]; The overview stream is processed to determine a user’s gaze direction. The virtual director uses rules for stream selection based on the user’s gaze direction.).
Regarding claim 42, Hafstad discloses everything claimed as applied above (see claim 31), in addition, Hafstad discloses, wherein the step of processing the acquired video data of the scene to detect an interaction state transition comprises the step of detecting a change in the looking direction (Gaze change; [0055]) of the at least one actor, or
comprises the step of performing speech detection ([0055]), and/or the step of monitoring an evolution of the body language of an actor ([0043], [0055]).
Regarding claim 45, Hafstad discloses everything claimed as applied above (see claim 31), in addition, Hafstad discloses, wherein the object of interest is at least one of another actor, an object (Lack of meaningful reactions in other objects; [0055]), a screen, or a camera.
Regarding claim 48, Hafstad discloses everything claimed as applied above (see claim 31), in addition, Hafstad discloses, further comprising the step of sending a command to a video mixer to show the video output data (The virtual director unit 203, 303 sends commands to the stream selector to show the video output data; [0049]), and/or wherein the video mixer is integrated in a system for selecting video output data (fig. 1).
Regarding claim 52, Hafstad discloses everything claimed as applied above (see claim 31), in addition, Hafstad discloses, A non-transitory computer program product comprising software which executed on one or more processing engines, performs the method of claim 31 (The software component of virtual director must be stored in some way to carry out the operation of the virtual director; [0050]).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 41, 44 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hafstad in view Ohno (US 2022/0075983 A1).
Regarding claim 41, Hafstad discloses everything claimed as applied above (see claim 40), in addition, Hafstad discloses, wherein the step of evaluating the looking direction of the at least one actor is performed by:
detecting at least one actor and the object of interest (The vision pipeline 202, 302 detects objects; [0043]),
estimating the looking direction of the at least one actor (The direction of the object’s gaze; [0054]),
.
However, Hafstad, fails to explicitly disclose the algorithm used to calculate the looking direction of the at least one actor. However, the examiner maintains that it was well known in the art to provide this, as taught by Ohno.
In a similar field of endeavor, Ohno discloses
performing 3D pose estimation to extract 3D body keypoints of the at least one actor (Face orientation information and 3D eye gaze information of the person 400; [0082]) and keypoints of the object of interest (Coordinate data of an eye gaze point on a predetermined target plane; [0082]),
estimating the looking direction of the at least one actor (Eye gaze vector or coordinate data of an eye gaze point; [0082]),
converting the looking direction of each actor to an angle in a reference coordinate system (Eye gaze vector; [0082]).
Hafstad teaches estimating the looking direction of an actor. Ohno teaches estimating the looking direction of a person by performing 3D pose information to extract 3D body keypoints of the person and keypoints of the object of interest and converting the looking direction of the person to an angle in a reference coordinate system. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to substitute the undisclosed algorithm for determining the looking direction of the actor in Hafstad with the algorithm for determining the looking direction of the person in Ohno to achieve the predictable result of accurately predicting where the person is looking.
Regarding claim 44, Hafstad and Ohno, the combination, discloses everything claimed as applied above (see claim 41), in addition, Hafstad discloses, wherein the step of estimating the looking direction is performed by at least one of facial landmark estimation to extract facial landmarks for at least one actor, head pose estimation, eye gaze analysis ([0054]-[0055]).
Claim(s) 43 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hafstad in view of Weitz et al. (US 2022/0148218 A1) hereinafter referenced as Weitz.
Regarding claim 43, Hafstad discloses everything claimed as applied above (see claim 42), however, Hafstad, fails to explicitly disclose the change in looking direction is detected by following an evolution of facial landmarks. However, the examiner maintains that it was well known in the art to provide this, as taught by Weitz.
In a similar field of endeavor, Weitz discloses wherein the change in the looking direction is detected by following an evolution of facial landmarks in each of the at least one actor ([0082]; The evolution of facial landmarks is the change in orientation of the eyes.).
Hafstad teaches determining a change in a gaze direction. Weitz teaches determining a change in gaze direction by analyzing a change in the orientation of the eyes. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to substitute the undisclosed method for determining the gaze direction change with analyzing a change in the orientation of the eyes to achieve the predictable result of accurately predicting where the person will be looking as disclosed in Weitz ([0203]).
Claim(s) 46-47, 49, 51 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hafstad in view of Official Notice.
Regarding claim 46, Hafstad discloses everything claimed as applied above (see claim 31), however, Hafstad fails to explicitly disclose calibrating the plurality of cameras to map the images viewed by each camera in each camera shot to a reference coordinate system. However, the examiner takes official notice of the fact that it was well known in the art before the effective filing date of the claimed invention (AIA ) to provide this.
Hafstad teaches a plurality of cameras used to capture a scene where different cameras and zoom levels are selected based on where people are looking. Calibrating a plurality of cameras capturing a scene to map images viewed by each camera to a reference coordinate system is well-known. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention (AIA ) to improve Hafstad by applying the technique of calibrating the plurality of cameras capturing the scene to map images viewed by each camera to a reference coordinate system to achieve the predictable result in generating an accurate estimation of where a person is looking.
Regarding claim 47, Hafstad discloses everything claimed as applied above (see claim 31), however, Hafstad fails to explicitly disclose that the camera shots have overlapping views to enable 3D reconstruction of the scene. However, the examiner takes official notice of the fact that it was well known in the art before the effective filing date of the claimed invention (AIA ) to provide this.
Hafstad teaches a plurality of cameras having different viewpoints capturing a scene. Providing an arrangement of cameras with overlapping views to enable 3D reconstruction is well-known. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention (AIA ) to improve Hafstad by applying the technique of providing the cameras with overlapping views to enable 3D reconstruction of the scene to achieve the predictable result in generating an accurate estimation of where a person is looking in the 3D environment.
Regarding claim 49, Hafstad discloses
A system for selecting video output data comprising
a plurality of cameras (100, 101; fig. 1)
a computer program product comprising software (Software component of virtual director; [0050]) which when executed on one or more processing engines, performs the method according to claim 31 to select the output video data,
a video mixer (Stream selector; fig. 1) configured to receive video streams from the plurality of cameras and to select the output video data based on the output of the computer program product.
However, Hafstad fails to explicitly disclose that the cameras are configured to enable a 3D reconstruction of a scene. However, the examiner takes official notice of the fact that it was well known in the art before the effective filing date of the claimed invention (AIA ) to provide this.
Hafstad teaches a plurality of cameras at different viewpoints capturing a scene. Using a plurality of cameras at different viewpoints to enable a 3D reconstruction of a scene is well-known. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention (AIA ) to improve Hafstad by applying the technique of providing the cameras with overlapping views to enable 3D reconstruction of the scene to achieve the predictable result in generating an accurate estimation of where a person is looking in the 3D environment.
Regarding claim 51, Hafstad discloses everything claimed as applied above (see claim 49), in addition, Hafstad discloses, wherein at least one camera is a [zoom] camera ([0039]).
However, Hafstad fails to explicitly disclose that at least one camera is a PTZ camera. However, the examiner takes official notice of the fact that it was well known in the art before the effective filing date of the claimed invention (AIA ) to provide this.
Hastad teaches a plurality of zoom cameras capturing a scene. PTZ cameras are well-known. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention (AIA ) to substitute the zoom cameras of Hafstad with PTZ cameras to achieve the predictable result of generating a number of views or shots larger than the number of cameras to reduce the cost of the system.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PAUL M BERARDESCA whose telephone number is (571)270-3579. The examiner can normally be reached Mon-Thurs 10-8, Fri 10-2.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sinh Tran can be reached at (571)272-7564. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
PAUL M. BERARDESCA
Examiner
Art Unit 2637
/PAUL M BERARDESCA/Primary Examiner, Art Unit 2637 7/15/2026