Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments with respect to claims 1-9 and 11-20 have been considered but are moot in view of a new ground of rejection necessitated by the amendments.
The examiner notes that Applicant has neither challenged nor mentioned the subject matter of the OFFICIAL NOTICE in Claims 11-12. Due to the applicants’ inadequate traversal of the examiner OFFICIAL NOTICE, the subject matter of the OFFICIAL NOTICE is taken to be applicants admitted prior art. See MPEP 2144.03(C) which recites “To adequately traverse such a finding, an applicant must specifically point out the supposed errors in the examiner’s action, which would include stating why the noticed fact is not considered to be common knowledge or well-known in the art” and “If applicant does not traverse the examiner’s assertion of official notice or applicant’s traverse is not adequate, the examiner should clearly indicate in the next Office action that the common knowledge or well-known in the art statement is taken to be admitted prior art because applicant either failed to traverse the examiner’s assertion of official notice or that the traverse was inadequate”. Clearly, Applicant did not state why the subject matter of the OFFICIAL NOTICE was not common knowledge or well known in the art.
The rejections for Claims 11-12 will be updated to reflect applicants admitted prior art.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-9 and 11-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 1, 15, and 20 recites “determining the scene data comprises generating a virtual representation that predicts a physical environment depicted in video data of a subset of the plurality of participants to use as the scene environment.
Scene data: is lighting / audio / perspective characteristics of a scene environment based on Applicants specification.
Scene environment: is a shared virtual background (e.g. camp fire / meeting room).
The specification clearly shows that scene data is what is used to display the selected scene environment from the plurality of candidate scene environments based on backgrounds of plural participants (Par.72, dark rooms / office environment). The plural backgrounds that most closely match one of the candidate scene environments is selected. To the examiner, lighting would define what is displayed in the shared background of Applicants Figure 6. How does scene data “predict” the physical environment in video of a subset of the plural participants? The limitation of “scene data generating a virtual representation that predicts a physical environment in video of a subset of the plural participants” renders the claim vague and indefinite.
Dependent Claims 2-9, 11-14, and 16-19 are further rejected as vague and indefinite as they do not remedy the issues in the Independent Claims.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 1-4, 7-9, 13-18, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Adcock (US 2022/0210374) in view of Roper (US 20230126108).
Regarding Claim 1, 15 and 20, Adcock teaches an immersive teleconferencing within a shared scene environment (Fig.1C: common virtual background), comprising: one or more processors; and one or more memory elements including instructions that when executed cause the one or more processors (Par.5) to: receive a plurality of streams for presentation at a teleconference (Par.5, Fig.1A, Fig.1B, 1st and 2nd video stream), wherein each of the plurality of streams represents a participant of a respective plurality of participants of the teleconference (Fig.1A, Fig.1B, User A, User B); determine scene data descriptive of a scene environment (Fig.1C:135, select common virtual background), wherein determining the scene data comprises generating a virtual representation (Fig.1C:135, selected and creates a virtual background) of a physical environment (Par.16, e.g virtual restaurant, conference room, park) to use as the scene environment (Par.16, 35, 37-39, based on at least user selection, calendar invite, agenda, time, or characteristics associated with the virtual session), the scene data comprising at least one of lighting characteristics (Par.65, lighting), {acoustic characteristics, or perspective characteristics of the scene environment}(no patentable weight given due to optional language); and for each of the plurality of participants of the teleconference: determine a position of the participant within the scene environment (Fig.1C:140, user A and B put in specific position); and based at least in part on the scene data, the position of the participant and at least one other participant within the scene environment, modify the stream that represents the participant such that the participant is depicted within the shared scene environment (Fig.1C, Par.59 and Par.42, places users within positions by adjusting the streams to place users into the common virtual background).
While Adcock does teach that the conference system can select common backgrounds based on various conditions of the virtual session (Par.16, 35, 37-39, based on at least user selection, calendar invite, agenda, time, or characteristics associated with the virtual session) Adcock does not explicitly teach generating/selecting a virtual representation that predicts a physical environment depicted in a video data of a subset of the plurality of participants to use as the scene environment.
Roper teaches that it is well known to generate/select a virtual representation that predicts a physical environment depicted in a video data of a subset of the plurality of participants to use as the scene environment (Par.87 and Par.89, selection of virtual background with similar characteristics to one or more other virtual backgrounds used by participants in the video conference). While this teaching in Roper is used to select a virtual background for a single user, this teaching/concept could obviously be incorporated into the selection of a common background as taught in Adcock.
Therefore, one of ordinary skill in the art would immediately recognize, before the effective filing date of the invention, that it would be obvious to modify Adcock’s immersive teleconferencing system wherein participants are presented within a common virtual scene environment based on characteristics of the virtual session, user selection, agenda, time --- with Roper’s teaching of generating or selecting a virtual background that predicts or matches physical environments depicted in video data of one or more participants to provide an enhanced system which provides more contextually relevant common virtual backgrounds leading to a more engaging, collaborative, and joined experience between the users. Further, both references address virtual background selection in video conferencing. The combination would have been straightforward for one of ordinary skill, as it merely augments Adcock’s background selection logic with Roper’s criteria, using known image analysis and selection techniques. This would predictably result in a system that can generate a common/shared virtual scene environment reflecting the actual environments of the users, as taught by Roper.
Regarding Claim 2 and 16, Adcock teaches the stream that represents the participant comprises at least video date that depicts the participant (Fig.1A).
Regarding Claim 3 and 17, Adcock teaches modifying the stream that represents the participant comprises modifying, by the computing system, the stream using one or more machine-learned models, wherein each of the machine-learned models are trained to process at least one of scene data or video data (Par.16 and Par.27, multiple video streams may be processed using machine learning or other techniques, such that the multiple video streams of the multiple users are positioned realistically within the simulated virtual environment).
Regarding Claims 4 and 18, Adcock teaches the one or more machine-learned models comprises a machine-learned semantic segmentation model trained to perform semantic segmentation tasks (Par.16, “The multiple video streams may be processed using machine learning or other techniques, such that the multiple video streams of the multiple users are positioned realistically within the simulated virtual environment”, the person is extracted or segmented from the video stream and placed into the virtual background);
the stream that represents the participant comprises the video data that depicts the participant (Fig.1A, Fig.1B); and wherein modifying the stream that represents the participant comprises segmenting, by the computing system, the video data of the stream that represents the participant into a foreground portion and a background portion (Fig.1C) using the machine-learned semantic segmentation model (Par.27 and Par.16, “The multiple video streams may be processed using machine learning or other techniques, such that the multiple video streams of the multiple users are positioned realistically within the simulated virtual environment”).
Regarding Claim 7, method of claim 2, Lanier teaches the stream that represents the participant comprises the video data that depicts the participant (Col.10:lines 13-51); {the scene data comprises the perspective characteristics of the scene environment, wherein the perspective characteristics indicate a perspective from which the scene environment is viewed; and wherein modifying the stream that represents the participant comprises: based at least in part on the perspective characteristics and the position of the participant within the scene environment, determining, by the computing system, that a portion of the participant that is visible from the perspective from which the scene environment is viewed is not depicted in the video data; generating, by the computing system, a predicted rendering of the portion of the participant; and applying, by the computing system, the predicted rendering of the portion of the participant to the video data}. (no patentable weight given, perspective was optional in claim 1 and given no weight).
Regarding Claim 8, method of claim 2, Lanier teaches the stream that represents the participant (Col.10:lines 13-51) {comprises the audio data that corresponds to the participant; the scene data comprises the acoustic characteristics of the scene environment; and wherein modifying the stream that represents the participant comprises modifying, by the computing system, the audio data based at least in part on the position of the participant within the scene environment relative to the acoustic characteristics of the scene environment}. (no patentable weight given, acoustics was optional in claim 1 and given no weight).
Regarding Claim 9, method of claim 1, Adcock teaches wherein receiving the plurality of streams further comprises receiving, by the computing system for each of the plurality of streams (Fig.1A, Fig.1B), scene environment data for the stream descriptive of lighting characteristics (Par.65, lighting), {acoustic characteristics} (no patentable weight given), or {perspective characteristics of the participant represented by the stream} (no patentable weight given); and wherein modifying the stream that represents the participant comprises: based at least in part on the scene data (Fig.1C, common virtual background), the position of the participant within the scene environment (Fig.1C), and the environment data for the stream (Par.65, lighting), modifying, by the computing system, the stream that represents the participant (Fig.1C, Par.42, Par.59, stream modified to place user in common virtual background).
Regarding 13, Adcock teaches modifying the stream that represents the participant comprises, based at least in part on the scene data and the position of the participant within the scene environment (Fig.1C: 145 and 150), modifying, by the computing system, the stream that represents the participant in relation to a position of another participant of the plurality of participants (Par.67, Par.72 and Fig.1C: 145 and 150, the stream is modified and what is presented/broadcast is the modified with individuals placed in specific positions (140)); and wherein the method further comprises broadcasting, by the computing system, the stream to a participant device respectively associated with the other participant (Fig.1C, transmitted to others at different locations, hence broadcast).
Regarding Claim 14, Adcock teaches generating, by the computing system, a shared stream that comprises the plurality of streams depicted within a virtualized representation of the scene environment based at least in part on the position of each of the plurality of participants within the scene environment (Par.67, Par.72 and Fig.1C: 145 and 150, the stream with common virtual background with individuals placed in specific positions (140)); and broadcasting, by the computing system, the shared stream to a plurality of participant devices respectively associated with the plurality of participants (Fig.1C, transmitted to others at different locations, hence broadcast).
Claims 5 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Adcock (US 2022/0210374) and Roper (US 20230126108) in further view of Lanier (US 11374988).
Regarding Claim 5 and 19, Adcock teaches the stream that represents the participant comprises the video data that depicts the participant (Fig.1A, Fig.1B, Fig.1C); the scene data describes the lighting characteristics of the scene environment (Par.65, visual effect-lighting), the lighting characteristics comprising a location and intensity of one or more light sources within the scene environment (Par.65, e.g. lighting of restaurant has light sources with varying location and intensity). Adcock already teachings placing people into a common virtual background and it would be obvious to adjust lighting based on position of the participants (Par.16 inserts images of people that appear realistically positioned); however Adcock and Roper does not explicitly teach wherein modifying the stream that represents the participant comprises: based at least in part on the scene data and the position of the participant, applying, by the computing system, a lighting correction to the video data that represents the participant based at least in part on the position of the participant within the scene environment relative to the one or more light sources.
Lanier teaches modifying the stream that represents the participant comprises: based at least in part on the scene data (Fig.5) and the position of the participant (Fig.5, seat position), applying, by the computing system, a lighting correction to the video data that represents the participant based at least in part on the position of the participant within the scene environment relative to the one or more light sources (Col.10:lines 13-51, Fig.5, Co.3:lines 46-52, lighting is adjusted relative streams 634(N) to blend/mask variations between individual streams, hence based, in part, on scene and position so that variations between streams are considered and masked/blended).
Therefore, to one of ordinary skill in the art before the effective filing date of the invention, it would have been obvious to modify Adcock and Ropers combined invention with the teachings of Lanier such that an enhanced system provides a video conference with users placed into a common/shared virtual background with the lighting adjusted to provide a realistic experience with the image looks blended or as natural as possible.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Adcock (US 2022/0210374) and Roper (US 20230126108) in further view of Locker (US 20230177884).
Regarding Claim 6, Adcock teaches the stream that represents the participant comprises the video data that depicts the participant (Fig.1A, 1B), however Adcock and Roper do not expressly teach but Locker teaches wherein the video data further depicts a gaze of the participant (Par.15, “the user looking away”, hence direction of gaze determined); and wherein modifying the stream that represents the participant comprises: determining, by the computing system, a direction of a gaze of the participant (Par.22, Par.15, Par.17, Par.31); determining, by the computing system, a gaze correction for the gaze of the participant (Par.22, Par.15, Par.17, Par.31). applying, by the computing system, the gaze correction to the video data to adjust the gaze of the participant depicted by the video data (Par.22, Par.15, Par.17, Par.31). Since Adcock teaches placement of a user into a common virtual background of a video conference, Lockers eye gaze correction in a video conference would naturally be based at least in part on the position of the participant within the scene environment and the gaze of the participant as it would occur at the spot the person is placed.
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date to modify Adcock and Ropers combined invention with the teachings of Locker such that an enhanced computing system is provided where eye gaze of a person within the virtual conference can be adjusted. This way the system can quickly identify a person that appears to be distracted and adjust the video such that the user appears laser focused providing a more professional setting with minimized distractions.
Claims 11-12 are rejected under 35 U.S.C. 103 as being unpatentable over Adcock (US 2022/0210374) and Roper (US 20230126108) in further view of in view of Applicants Admitted Prior Art.
Regarding Claim 11 and 12, Adcock teaches determining, by the computing system, the scene data descriptive of the scene environment comprises: determining, by the computing system, a participant scene environment for the plurality of streams (Par.16, virtual environment can be a park, restaurant, conference room, etc.) and Roper teaches that it is well known to generate/select a virtual representation that predicts a physical environment depicted in a video data of a subset of the plurality of participants to use as the scene environment (Par.87 and Par.89, selection of virtual background with similar characteristics to one or more other virtual backgrounds used by participants in the video conference). While this teaching in Roper is used to select a virtual background for a single user, this teaching/concept can obviously be incorporated into the selection of a common background as taught in Adcock.
And further, Applicants Admitted Prior Art teaches determining, by the computing system, a plurality of participant scene environments for the plurality of streams; and based at least in part on the plurality of participant scene environments, selecting, by the computing system, the scene environment from a plurality of candidate scene environments in view of Applicants silence to the Official Notice in the prior rejection.
Applicants Admitted Prior Art teaches that it is well known in the art for a video conference computing system to select one of a plurality of backgrounds/scene environments. Therefore, it would have been obvious before the effective filing date of the invention to modify Adcock and Roper using a shared background for conference attendees such that an enhanced computing system is provided which can select one of a plurality of backgrounds/scene environments allowing selection of a default background/scene environment upon startup of a meeting but can also allow for customization of meetings where a different background/scene environment can be selected based on the type of meeting (i.e. casual/business) or mood of the person hosting the meeting.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WESLEY LEO KIM whose telephone number is (571)272-7867. The examiner can normally be reached 9-5:30 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WESLEY L KIM/Supervisory Patent Examiner, Art Unit 2648