DETAILED ACTION
This action is in response to the application filed 11/13/2024. Claims 1 – 20 are pending and have
been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 4 – 9 and 15 - 18 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 4 recites the limitations “the multiple cameras” and “the multiple microphones”. Claim 4 is dependent on Claim 1, which recites “one or more cameras” and “one or more microphones”. There is insufficient antecedent basis for “the multiple cameras” and “the multiple microphones” because Claim 1 does not previously introduce explicitly “multiple cameras” or “multiple microphones”.
Claims 5 – 9 are dependent, directly or indirectly, on Claim 4 and therefore are similarly indefinite because they incorporate the indefinite limitations of Claim 4.
Claim 15 recites the limitation “the multiple microphones”. Claim 15 is dependent on Claim 14, which recites “one or more microphones”. There is insufficient antecedent basis for “the multiple microphones” because Claim 14 does not previously introduce explicitly “multiple microphones”.
Claim 16 is dependent, directly or indirectly, on Claim 15 and therefore is similarly indefinite because it incorporates the indefinite limitations of Claim 15.
Claim 17 recites the limitation “the cameras”. Claim 17 is dependent on Claim 14, which recites “one or more cameras”. There is insufficient antecedent basis for “the cameras” because Claim 14 does not previously introduce explicitly “the cameras”.
Claim 18 is dependent, directly or indirectly, on Claim 17 and therefore is similarly indefinite because it incorporates the indefinite limitations of Claim 17.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 2, 10, 13, 14, 19 and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Fitzpatrick et al. (U.S. Patent No. 11,924,258, hereinafter “Fitzpatrick”).
Regarding Claim 1, Fitzpatrick teaches
A collaborative operating system of audiovisual peripheral devices (see Fitzpatrick Column 2, lines 53 – 56, the system comprises at least one cooperation application to configure the capture device and the control device to cooperate during the performance of a video processing operation), comprising:
an audiovisual console (see Fitzpatrick Abstract, control device), wherein the audiovisual console includes an audiovisual processing unit (see Fitzpatrick Column 12, lines 65 – 67 and Column 13, line 1, the control device 5 comprises many equivalent components, such as a processor 53), a microcontroller used to operate collaboration of audiovisual peripherals (see Fitzpatrick Column 13, lines 4 – 17, The control device 5 is configured to allow a user to download, via the telecommunication module 52, both a video conferencing program 62 and a video control program 61 to the memory where each can be simultaneously executed by the control device 5. The control device 5 also comprises its own pairing module 56, which is coupled to a computer port 58, such as a USB port. Again, this can be used to connect the control device 5 to the capture device 10 via the wired local connection 6, as described above, that allows transfer of both power and data between the two devices 5, 10), at least one connection interface (see Fitzpatrick Column 11, lines 39 – 47, pairing the capture device 10 with the control device 5 so their respective shortcomings can be synergistically overcome. To achieve this, the system 1 further comprises a local connection 6 via which the capture device 10 and the control device 5 can be communicatively coupled to one another. In alternative embodiments, they may be communicatively coupled in other ways—for example, a wireless connection can take the place of the local wired connection 6), and a data-processing unit (see Fitzpatrick Column 14, lines 58 – 61, A video stream, captured and processed by the capture device 10 can be transmitted via the local connection 6 to the control device 5 for use in a video conference, and Column 9, lines 9 – 13, The app 21, and the video control program 61 are each specific instances of a cooperation application, as will be described below, that allow the mobile device and computing device to cooperate with one another in the performance of a video processing operation) used to process data transmitted via the at least one connection interface (see Fitzpatrick Column 9, lines 32 – 39, the at least one cooperation program allows a capture device and a control device to be configured to operate together, rather than as independent standalone devices. Thus, their respective shortcomings can be synergistically offset by one another, and they can share the burden of tasks required for video processing, including the potential enhancement of a video stream—ideally originating from the capture device, and Figure 2, displaying components of the control device and the connection to the capture device);
wherein the audiovisual console connects with at least one peripheral device via the at least one connection interface based on a communication protocol (see Fitzpatrick Column 10, lines 44 – 51, The communications network 2 interconnects the components of the system 1, as well as the components that interact with the system 1. In various embodiments the network may be embodied by a wired and/or wireless local area network (LAN), peer-to-peer wireless connections (e.g. using at least one of Bluetooth and direct Wi-Fi), a wide area network (WAN) such as the Internet, or a combination of these), and detects one or more cameras and one or more microphones of the at least one peripheral device (see Fitzpatrick Column 6, lines 57 – 65, this allow the cooperation application to efficiently split the processing burden between the capture device and the control device depending on their likely video processing capabilities. Moreover, the executed at least one cooperation application ideally determines the respective technical capabilities of the capture device and the control device and, in dependence on the determined technical capabilities, determines the split of video processing tasks between the capture device and the control device, Column 7, lines 11 – 17, determining the relative technical capabilities of each device, and then allocating tasks to the two devices in dependence on their relative technical capabilities. To this end, the method may further comprise exchanging technical capability data between the capture device and the control device, Column 12, lines 19 – 24, The capture device 10 also has an audio input/output 19, in the form of a microphone array, and internal speakers. The audio output typically also includes a stereo output interface—such a headphone jack, or wireless audio transmitter—via which stereo sound signals can be transmitted to stereo headphones or other stereo sound generation means, and Figure 2, capture device and control device both include audio input and output);
wherein, after a permission for accessing the one or more cameras and the one or more microphones is obtained (see Fitzpatrick Column 13, lines 53 – 65, The pairing routine may comprise an authorisation process to ensure that a user is authorised to use both the control device 5 and the capture device 10 in conjunction with one another. This is less important for the current embodiment in which a direct wired location connection 6 is used, than in alternative embodiments in which a wireless connection takes the place of the wired location connection. The authorisation process may comprise a key or code exchange between the control device 5 and the capture device 10 in a way that maximises the likelihood that both devices are under the control of the same authorised user. The authorisation process may comprise outputting a code or signal at one device, for receipt at the other device,), the audiovisual console receives a video and an audio from the at least one peripheral device via the at least one connection interface based on the communication protocol (see Fitzpatrick Column 14, lines 56 – 62, The capture device 10 can thus be effectively utilised as an efficient high-quality independent webcam, and is ideally is secured to a position adjacent to the control device 5. A video stream, captured and processed by the capture device 10 can be transmitted via the local connection 6 to the control device 5 for use in a video conference, as governed by the video conferencing program 62, Column 12, lines 19 – 24, The capture device 10 also has an audio input/output 19, in the form of a microphone array, and internal speakers. The audio output typically also includes a stereo output interface—such a headphone jack, or wireless audio transmitter—via which stereo sound signals can be transmitted to stereo headphones or other stereo sound generation means, Column 14, lines 6 – 8, Another example is via outputting an audio signal from a speaker of one device, to be detected by a microphone by the other, and Figure 2, capture device and control device both include audio input and output, Column 18, lines 21 – 27, By transmitting technical capability data between the devices for comparison, an efficient split of tasks for conducting video processing (including audio processing) can be achieved, which is particularly important for real-time video communications such as video streaming and/or video conferencing); wherein, after the data-processing unit processes the video and the audio, the microcontroller generates audiovisual data provided to the audiovisual processing unit (see Fitzpatrick Column 3, lines 24 – 29, the use of at least one cooperation program for pairing allows video from the camera of the capture device to be fed, via the respective pairing modules, to the control device for use in a video processing operation, such as for recording, and/or real-time video communications—such as a video conferencing session); wherein, after the audiovisual processing unit processes the audiovisual data, the audiovisual console uses a display to display the video, and uses a speaker to play the audio (see Fitzpatrick Column 3, lines 66 – 7 and Column 4, lines 1 – 13, the control device comprises a display unit, and the cooperation application configures the control device to display on the display unit a user interface (UI). Preferably, the UI has at least one UI element that is configured to receive a user input so as to: change settings of the camera of the capture device, such as brightness, contrast, depth of field, bokeh effect and/or image resolution; specify video processing tasks to be performed at the capture device and/or control device; display video, on the display unit, of video generated by the capture device; start video generation by the camera of the capture device; and/or stop video generation by the camera of the capture device, Column 4, lines 35 – 38, Preferably, the control device further comprises a display unit configured to display the video streams transmitted and received by the control device, Column 12, lines 19 – 20, capture device includes microphone array, and Column 2, lines 29 – 33, the control device comprises an audio interface for generating video stream audio signals).
Regarding Claim 2, Fitzpatrick teaches
The collaborative operating system according to claim 1, wherein, when the audiovisual console and the at least one peripheral device are located within a same local area network or a same area, the audiovisual console and the at least one peripheral device are interconnected by a wireless local area network or a Bluetooth communication protocol (see Fitzpatrick Column 10, lines 44 – 51, The communications network 2 interconnects the components of the system 1, as well as the components that interact with the system 1. In various embodiments the network may be embodied by a wired and/or wireless local area network (LAN), peer-to-peer wireless connections (e.g. using at least one of Bluetooth and direct Wi-Fi), a wide area network (WAN) such as the Internet, or a combination of these, Column 11, lines 42 – 47, the system 1 further comprises a local connection 6 via which the capture device 10 and the control device 5 can be communicatively coupled to one another. In alternative embodiments, they may be communicatively coupled in other ways—for example, a wireless connection can take the place of the local wired connection 6).
Regarding Claim 10, Fitzpatrick teaches
The collaborative operating system according to claim 1, wherein a user uses identification data to log in a server program of the audiovisual console and an audiovisual program of the at least one peripheral device (see Fitzpatrick Column 13, lines 53 – 65, The pairing routine may comprise an authorisation process to ensure that a user is authorised to use both the control device 5 and the capture device 10 in conjunction with one another. The authorisation process may comprise a key or code exchange between the control device 5 and the capture device 10 in a way that maximises the likelihood that both devices are under the control of the same authorised user. The authorisation process may comprise outputting a code or signal at one device, for receipt at the other device, Column 10, line 25, the video conferencing server 3, Column 4, lines 22 – 25, the code exchange comprises generating and outputting as a video or audio signal, a code at one of the capture device or control device, and receiving and inputting that code at the other of the capture device or control device), so that a connection is established between the audiovisual console and the at least one peripheral device (see Fitzpatrick Column 4, lines 26 – 28, the control device further comprises a networking module configured to establish a connection, via a communications network), and the audiovisual console obtains the permission for accessing the one or more cameras and the one or more microphones of the at least one peripheral device (see Fitzpatrick Column 13, lines 53 – 65, The pairing routine may comprise an authorisation process to ensure that a user is authorised to use both the control device 5 and the capture device 10 (comprising peripheral camera and microphone array) in conjunction with one another. The authorisation process may comprise a key or code exchange between the control device 5 and the capture device 10 in a way that maximises the likelihood that both devices are under the control of the same authorised user. The authorisation process may comprise outputting a code or signal at one device, for receipt at the other device).
Regarding Claim 13, Fitzpatrick teaches
The collaborative operating system according to claim 1, wherein the audiovisual console and the at least one peripheral device form a first conference terminal (see Fitzpatrick Column 14, lines 56 – 62, The capture device 10 can thus be effectively utilised as an efficient high-quality independent webcam, and is ideally is secured to a position adjacent to the control device 5. A video stream, captured and processed by the capture device 10 can be transmitted via the local connection 6 to the control device 5 for use in a video conference, as governed by the video conferencing program 62), and the first conference terminal is configured to connect with a second conference terminal, so as to establish a conference session (see Fitzpatrick Column 2, lines 24 – 50, Preferably, the system is a video-conferencing system. The system may comprise at least one of: a control device and a capture device. Preferably, the control device comprises at least one of: a networking module; a display unit for displaying video streams; an audio interface for generating video stream audio signals; and a pairing module for communicatively pairing the control device with the capture device. Preferably, the networking module of the control device is configured to establish a connection, via a communications network, with at least one recipient of a video stream transmitted by the control device. The at least one recipient of the video stream transmitted by the control device may be another participant of a video-conferencing session. Preferably, the display device of the control device is configured to display video streams, such as those of the video-conferencing session. Preferably, the capture device comprises at least one of: a screen, such an electronic touch-sensitive screen; a sensor set, including a camera; and a pairing module for communicatively pairing the capture device with the control device; and a telecommunications module); wherein the audiovisual console of the first conference terminal uses the display to display the video received from the at least one peripheral device, and uses the speaker to play the audio received from the at least one peripheral device, so that a video conference is in operation (see Fitzpatrick Column 4, lines 35 – 38, Preferably, the control device further comprises a display unit configured to display the video streams transmitted and received by the control device, Column 12, lines 19 – 24, The capture device 10 also has an audio input/output 19, in the form of a microphone array, and internal speakers. The audio output typically also includes a stereo output interface—such a headphone jack, or wireless audio transmitter—via which stereo sound signals can be transmitted to stereo headphones or other stereo sound generation means, Figure 2, capture device and control device both include audio input and output).
Regarding Claims 14, 19 and 20, they are rejected similarly as Claims 1, 10 and 13, respectively. The method can be found in Fitzpatrick (Abstract, method).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 3 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Fitzpatrick et al. (U.S. Patent No. 11,924,258, hereinafter “Fitzpatrick”) in view of Bettenbuk et al. (U.S. Pub. No. 2017/0359187, hereinafter “Bettenbuk”).
Regarding Claim 3, Fitzpatrick teaches
The collaborative operating system according to claim 1, wherein, when the audiovisual console and the at least one peripheral device are not located within a same local area network (see Fitzpatrick Column 10, lines 44 – 51, The communications network 2 interconnects the components of the system 1, as well as the components that interact with the system 1. In various embodiments the network may be embodied by a wired and/or wireless local area network (LAN), peer-to-peer wireless connections (e.g. using at least one of Bluetooth and direct Wi-Fi), a wide area network (WAN) such as the Internet, or a combination of these, Column 11, lines 42 – 47, the system 1 further comprises a local connection 6 via which the capture device 10 and the control device 5 can be communicatively coupled to one another. In alternative embodiments, they may be communicatively coupled in other ways—for example, a wireless connection can take the place of the local wired connection 6),
Fitzpatrick does not expressively teach
a connection between the audiovisual console and the at least one peripheral device is established by a software program running an Interactive Connectivity Establishment (ICE) communication protocol or a Web Real-Time Communication (WebRTC) protocol
However, Bettenbuk teaches
a connection between the audiovisual console and the at least one peripheral device is established by a software program running an Interactive Connectivity Establishment (ICE) communication protocol or a Web Real-Time Communication (WebRTC) protocol (see Bettenbuk Paragraph [0021], The Web Real-Time communication (WebRTC) framework provides the protocol building blocks to support direct, interactive, real-time communication using audio, video, collaboration, etc., between two peers' web-browsers. WebRTC uses the Real-time Transport Protocol (RTP) (RFC3550) as its media transport protocol. RTP provides a framework for delivery of audio and video teleconferencing data and other real-time media applications).
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the teaching of a system in which an audiovisual console connects to one or more peripheral devices through a communication interface, detects components of the peripheral device, obtains permission to access them, receives audio and video data, and generates audiovisual output to present as video on a display and audio to play on the speakers (as taught in Fitzpatrick), with connection between devices being established using a software program that operates over a WebRTC protocol (as taught in Bettenbuk), the motivation being to enable real-time videoconferencing with low latency connections and very smart logic in terms of handling receiver feedbacks (e.g. packet loss) (see Bettenbuk Paragraph [0003]).
Regarding Claim 11, Fitzpatrick in view of Bettenbuk teaches
The collaborative operating system according to claim 10, wherein the audiovisual console executes a video conference program or a software program of a Web Real-Time Communication (WebRTC) protocol (see Bettenbuk Paragraph [0021], The Web Real-Time communication (WebRTC) framework provides the protocol building blocks to support direct, interactive, real-time communication using audio, video, collaboration, etc., between two peers' web-browsers. WebRTC uses the Real-time Transport Protocol (RTP) (RFC3550) as its media transport protocol. RTP provides a framework for delivery of audio and video teleconferencing data and other real-time media applications).
Claims 4 and 5 are rejected under 35 U.S.C. 103 as being unpatentable over Fitzpatrick et al. (U.S. Patent No. 11,924,258, hereinafter “Fitzpatrick”) in view of Cutler et al. (U.S. Pub. No. 2004/0001137, hereinafter “Cutler”).
Regarding Claim 4, Fitzpatrick teaches
The collaborative operating system according to claim 1, wherein the at least one connection interface is a wireless or wired communication interface that is used to receive the video and audio (see Fitzpatrick Column 13, lines 4 – 19, The control device 5 is configured to allow a user to download, via the telecommunication module 52, both a video conferencing program 62 and a video control program 61 to the memory where each can be simultaneously executed by the control device 5. These programs provide some of the functionality of the video processing system 1—in this case performed by the control device 5. The control device 5 also comprises its own pairing module 56, and power supply unit 57, which are each coupled to a computer port 58, such as a USB port. Again, this can be used to connect the control device 5 to the capture device 10 via the wired local connection 6, as described above, that allows transfer of both power and data between the two devices 5, 10. The control device 5 may further comprise or be connectable with an auxiliary control device, such as a keyboard or another peripheral, Column 12, lines 65 – 67 and Column 13, line 1, the control device 5 comprises many equivalent components, such as a telecommunication module 52 for interfacing with the network 2, a processor 53, an audio i/o module 55 and memory 54, Column 9, lines 32 – 39, the at least one cooperation program allows a capture device and a control device to be configured to operate together, rather than as independent standalone devices. Thus, their respective shortcomings can be synergistically offset by one another, and they can share the burden of tasks required for video processing, including the potential enhancement of a video stream—ideally originating from the capture device, and Figure 2, displaying components of the control device and the connection to the capture device)
Fitzpatrick does not expressively teach
The video captured through lenses of the multiple cameras and the audio received by the multiple microphones of the at least one peripheral device .
However, Cutler teaches
The video captured through lenses of the multiple cameras and the audio received by the multiple microphones of the at least one peripheral device (see Cutler Abstract, An omni-directional camera (a 360 degree camera) is proposed with an integrated microphone array. The primary application for such a camera is videoconferencing and meeting recording, and the device is designed to be placed on a meeting room table. The microphone array is in a planar configuration, and the microphones are located as close to the desktop as possible to eliminate sound reflections from the table. The camera is connected to the microphone array base with a thin cylindrical rod, which is acoustically invisible to the microphone array for the frequency range [50-4000] Hz. This provides a direct path from the person talking to all of the microphones in the array, and can therefore be used for sound source localization (determining the location of the talker) and beam-forming (improving the sound quality of the talker by filtering only sound from a particular direction), Paragraph [0061], If a multi-sensor camera configuration is used, a plurality of camera or video sensors can be employed, Paragraph [0080], an omni-directional camera is used that employs multiple video sensors to achieve 360 degree camera coverage. Alternately, in another embodiment of the invention, an omni-directional camera that employs one video sensor and a hyperbolic lens that captures light from 360 degrees to achieve panoramic coverage is used. Furthermore, either of these cameras may be used by themselves, elevated on the acoustically transparent cylindrical rod, to provide a frontal view of the meeting participants, Paragraph [0086], As mentioned previously, the video and audio signals can be broadcast to another video conferencing site or the Internet).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a system in which an audiovisual console connects to one or more peripheral devices through a communication interface, detects components of the peripheral device, obtains permission to access them, receives audio and video data, and generates audiovisual output to present as video on a display and audio to play on the speakers (as taught in Fitzpatrick), with capturing video through multiple cameras and audio through multiple microphones (as taught in Cutler), the motivation being to enable a system to have a selection of cameras and microphones to use during a conference to provide improved image quality, optimal viewing angles, enhanced audio capture, and more accurate sound localization (see Cutler Paragraphs [0004] – [0009]).
Regarding Claim 5, Fitzpatrick in view of Cutler teaches
The collaborative operating system according to claim 4, wherein the multiple microphones form a microphone array, so as to receive the audio from the microphones at different positions and perform multi-channel beamforming for tracking a sound source (see Cutler Abstract, An omni-directional camera (a 360 degree camera) is proposed with an integrated microphone array. The primary application for such a camera is videoconferencing and meeting recording, and the device is designed to be placed on a meeting room table. The microphone array is in a planar configuration, and the microphones are located as close to the desktop as possible to eliminate sound reflections from the table. The camera is connected to the microphone array base with a thin cylindrical rod, which is acoustically invisible to the microphone array for the frequency range [50-4000] Hz. This provides a direct path from the person talking to all of the microphones in the array, and can therefore be used for sound source localization (determining the location of the talker) and beam-forming (improving the sound quality of the talker by filtering only sound from a particular direction)).
Claims 6 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Fitzpatrick et al. (U.S. Patent No. 11,924,258, hereinafter “Fitzpatrick”) in view of Cutler et al. (U.S. Pub. No. 2004/0001137, hereinafter “Cutler”) and Lu et al. (U.S. Pub. No. 2017/0070814, hereinafter “Lu”).
Regarding Claim 6, Fitzpatrick in view of Cutler teaches all the limitations of claim 5, but does not expressively teach
The collaborative operating system according to claim 5, wherein the microcontroller of the audiovisual console obtains information of a volume of each of the microphones, and performs positioning on the multiple microphones according to audio magnitude of each of the microphones.
However, Lu teaches
The collaborative operating system according to claim 5, wherein the microcontroller of the audiovisual console obtains information of a volume of each of the microphones, and performs positioning on the multiple microphones according to audio magnitude of each of the microphones (see Lu Paragraph [0029], the directions of sound sources are from the front, back, left, right, top, and bottom surfaces of the device, and can be determined by amplitude and phase differences of microphone signals with proper microphone positioning. The sound source separation separates the sound coming from different directions from a mix of sources in microphone signals and identifies the direction of the sound sources. In some microphone placement implementations, sound source separation can be further performed using blind source separation (BSS), independent component analysis (ICA), and beamforming (BF) technologies).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a system in which an audiovisual console connects to one or more peripheral devices through a communication interface, detects multiple cameras and multiple microphones of the peripheral device, obtains permission to access them, receives audio and video data, and generates audiovisual output to present as video on a display and audio to play on the speakers (as taught in Fitzpatrick in view of Cutler), with measuring audio level from each of multiple microphones and determining the location of the audio source based on the relative audio magnitudes detected by the microphones (as taught in Lu), the motivation being to enable a system to prioritize and beamform the microphones closest to the speaker, reducing background noise and improving speech clarity (see Lu Abstract).
Regarding Claim 15, Fitzpatrick in view of Cutler and Lu teaches
The operating method according to claim 14, wherein the multiple microphones form a microphone array (see Cutler Abstract, An omni-directional camera (a 360 degree camera) is proposed with an integrated microphone array), so as to receive the audio from the microphones at different positions and perform multi-channel beamforming for tracking a sound source (see Cutler Abstract, An omni-directional camera (a 360 degree camera) is proposed with an integrated microphone array. The primary application for such a camera is videoconferencing and meeting recording, and the device is designed to be placed on a meeting room table. The microphone array is in a planar configuration, and the microphones are located as close to the desktop as possible to eliminate sound reflections from the table. The camera is connected to the microphone array base with a thin cylindrical rod, which is acoustically invisible to the microphone array for the frequency range [50-4000] Hz. This provides a direct path from the person talking to all of the microphones in the array, and can therefore be used for sound source localization (determining the location of the talker) and beam-forming (improving the sound quality of the talker by filtering only sound from a particular direction)); wherein, when the microcontroller of the audiovisual console obtains information of a volume of each of the microphones (see Lu Paragraph [0029], the directions of sound sources are from the front, back, left, right, top, and bottom surfaces of the device, and can be determined by amplitude and phase differences of microphone signals with proper microphone positioning), the microcontroller performs positioning on the multiple microphones according to audio magnitude of each of the microphones (see Lu Paragraph [0029], the directions of sound sources are from the front, back, left, right, top, and bottom surfaces of the device, and can be determined by amplitude and phase differences of microphone signals with proper microphone positioning. The sound source separation separates the sound coming from different directions from a mix of sources in microphone signals and identifies the direction of the sound sources. In some microphone placement implementations, sound source separation can be further performed using blind source separation (BSS), independent component analysis (ICA), and beamforming (BF) technologies).
Claims 7 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Fitzpatrick et al. (U.S. Patent No. 11,924,258, hereinafter “Fitzpatrick”) in view of Cutler et al. (U.S. Pub. No. 2004/0001137, hereinafter “Cutler”), Lu et al. (U.S. Pub. No. 2017/0070814, hereinafter “Lu”) and Elko (U.S. Patent No. 6,041,127).
Regarding Claim 7, Fitzpatrick in view of Cutler and Lu teaches
The collaborative operating system according to claim 6, wherein, after the positioning of the multiple microphones is completed, gains of the multiple microphones are adjusted according to a location of a user (see Cutler Paragraph [0021], The audio input can be also be used for various purposes. For instance, the audio can be used for sound source localization, so that the audio can be optimized for the speaker's direction at any given time. Additionally, a beam forming module can be used in the computer to improve the beam shape of the audio thereby further improving filtering of audio from a given direction. A noise reduction and automatic gain control module can also be used to improve the signal to noise ratio by reducing the noise and adjusting the gain to better capture the audio signals from a speaker, as opposed to the background noise of the room).
Fitzpatrick in view of Cutler and Lu does not expressively teach
positions, and orientations, of the multiple microphones are adjusted according to a location of a user.
However, Elko teaches
positions, and orientations, of the multiple microphones are adjusted according to a location of a user (see Elko Column 1, lines 44 – 51, The present invention provides a microphone array having a steerable response pattern, wherein the microphone array comprises a plurality of individual pressure-sensitive omnidirectional microphones and a processor adapted to compute difference signals between the pairs of the individual microphone output signals and to selectively combine these difference signals so as to produce a response pattern having an adjustable orientation of maximum reception, Column 17, lines 20 – 22, program may be advantageously written to allow for eight general first-order beam outputs that can be steered to any direction in 4.pi. space, Column 6, lines 8 – 13, FIG. 3 shows an illustrative computed output of a 30.degree. synthesized dipole microphone rotated by 30.degree., derived from four omnidirectional microphones arranged as illustratively shown in FIG. 2. The element spacing d is 2.0 cm and the frequency is 1 kHz. FIG. 4 shows an illustrative frequency response in the direction along the dipole axis for a 30.degree).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a system in which an audiovisual console connects to one or more peripheral devices through a communication interface, detects multiple cameras and multiple microphones of the peripheral device, obtains permission to access them, receives audio and video data, and generates audiovisual output to present as video on a display and audio to play on the speakers, determines the location of a user based on the relative audio levels detected by the multiple microphones, and adjusts the gains of the microphones according to the user’s location (as taught in Fitzpatrick in view of Cutler and Lu), with adjusting positions and orientations of microphones according to the location of a user (as taught in Elko), the motivation being to enable a system that electronically steers the microphone array toward the desired speaker, thus improving audio capture, without requiring the microphones to be physically repositioned (see Elko Abstract).
Regarding Claim 16, it is rejected similarly as Claim 7. The method can be found in Fitzpatrick (Abstract, method).
Claims 8, 9, 17 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Fitzpatrick et al. (U.S. Patent No. 11,924,258, hereinafter “Fitzpatrick”) in view of Cutler et al. (U.S. Pub. No. 2004/0001137, hereinafter “Cutler”) and Moshfeghi (U.S. Pub. No. 2017/0053419).
Regarding Claim 8, Fitzpatrick in view of Cutler teaches all the limitations of claim 4, but does not expressively teach
The collaborative operating system according to claim 4, wherein the microcontroller retrieves images captured through the lenses of the cameras, and an image-processing technology is used to compare the images of the video captured through each of the lenses, so as to determine a position of each of the cameras.
However, Moshfeghi teaches
The collaborative operating system according to claim 4, wherein the microcontroller retrieves images captured through the lenses of the cameras (see Moshfeghi Paragraph [0029], In some embodiments, process 200 is performed by a positioning server such as the positioning server 160 described by reference to FIG. 1, above. The process, in phase one, detects and positions (at 205) an object such as a human by using images received from a group of still image and/or video cameras), and an image-processing technology is used to compare the images of the video captured through each of the lenses, so as to determine a position of each of the cameras (see Moshfeghi Paragraph [0029], In some embodiments, process 200 is performed by a positioning server such as the positioning server 160 described by reference to FIG. 1, above. The process, in phase one, detects and positions (at 205) an object such as a human by using images received from a group of still image and/or video cameras. The term image is used herein to refer to still as well as video images. In phase one, multiple image feeds (still or video images) from cameras, such as camera 1 140 and camera 2 145 shown in FIG. 1, are used to identify and estimate the location of the objects, Paragraph [0030], Image processing, pattern recognition, face recognition, and feature detection algorithms are used by the positioning server to first identify a person or an object, such as object 1 105, and then estimate its location. In some embodiments, the object is identified, e.g., by comparing/correlating the images taken by the cameras with a set of images stored in a database, Paragraph [0032], Once an object is identified in multiple still/video image feeds, the two-dimensional position of that object (e.g., object 1 105 in FIG. 1) is extracted from the multiple video feeds of cameras (e.g., camera 1 140 and camera 2 145 in FIG. 1). Given the position of camera 1 140 and camera 2 145 (e.g., their x, y, z, coordinates in a coordinate system used by the positioning server, latitude/longitude/height, etc.), their orientation, and the extracted two-dimensional position of the object in the video/image snapshots, a 3-dimensional location for the object is estimated by using techniques such as geometrical correlation, geometrical triangulation, and triangular mapping).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a system in which an audiovisual console connects to one or more peripheral devices through a communication interface, detects multiple cameras and multiple microphones of the peripheral device, obtains permission to access them, receives audio and video data, and generates audiovisual output to present as video on a display and audio to play on the speakers (as taught in Fitzpatrick in view of Cutler), with comparing images from a video to determine positions of cameras (as taught in Moshfeghi), the motivation being to enable a system to accurately identify, position, and track objects in an environment (see Moshfeghi Paragraph [0006]).
Regarding Claim 9, Fitzpatrick in view of Cutler and Moshfeghi teaches
The collaborative operating system according to claim 8, wherein, in the audiovisual console, a location of a user is determined by performing three-dimensional detection on the multiple images captured through the lenses of the multiple cameras (see Moshfeghi Paragraph [0029], In some embodiments, process 200 is performed by a positioning server such as the positioning server 160 described by reference to FIG. 1, above. The process, in phase one, detects and positions (at 205) an object such as a human by using images received from a group of still image and/or video cameras. The term image is used herein to refer to still as well as video images. In phase one, multiple image feeds (still or video images) from cameras, such as camera 1 140 and camera 2 145 shown in FIG. 1, are used to identify and estimate the location of the objects, Paragraph [0030], Image processing, pattern recognition, face recognition, and feature detection algorithms are used by the positioning server to first identify a person or an object, such as object 1 105, and then estimate its location. Methods, such as correlating images taken at different angles from multiple cameras, are used for this purpose. In some embodiments, the object is identified, e.g., by comparing/correlating the images taken by the cameras with a set of images stored in a database, Paragraph [0032], Once an object is identified in multiple still/video image feeds, the two-dimensional position of that object (e.g., object 1 105 in FIG. 1) is extracted from the multiple video feeds of cameras (e.g., camera 1 140 and camera 2 145 in FIG. 1). Given the position of camera 1 140 and camera 2 145 (e.g., their x, y, z, coordinates in a coordinate system used by the positioning server, latitude/longitude/height, etc.), their orientation, and the extracted two-dimensional position of the object in the video/image snapshots, a 3-dimensional location for the object is estimated by using techniques such as geometrical correlation, geometrical triangulation, and triangular mapping. For instance, geometrical triangulation determines the location of an object by measuring angles to the object from at least two known points. The location of the object is then determined as the third point of a triangle with one known side and two known angles).
Regarding Claims 17 - 18, they are rejected similarly as Claims 8 - 9, respectively. The method can be found in Fitzpatrick (Abstract, method).
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Fitzpatrick et al. (U.S. Patent No. 11,924,258, hereinafter “Fitzpatrick”) in view of Vaughan et al. (U.S. Pub. No. 2024/0118744, hereinafter “Vaughan”).
Regarding Claim 12, Fitzpatrick teaches all the limitations of claim 10, but does not expressively teach
The collaborative operating system according to claim 10, wherein the audiovisual console includes a regulation program that is used to regulate the video and the audio, determine positions of the one or more microphones and the one or more cameras of the at least one peripheral device, determine a position of the speaker connected with the audiovisual console, regulate post-processing and a volume of the speaker, and regulate brightness and chrominance of a picture displayed on the display.
However, Vaughan teaches
The collaborative operating system according to claim 10, wherein the audiovisual console includes a regulation program that is used to regulate the video and the audio (see Vaughan Paragraph [0119], In conjunction with touch screen 212, display controller 256, contact/motion module 230, graphics module 232, audio circuitry 210, speaker 211, RF circuitry 208, text input module 234, e-mail client module 240, and browser module 247, online video module 255 includes instructions that allow the user to access, browse, receive (e.g., by streaming and/or download), Paragraph [0265], The user may also initiate various functions which involve state changes at the device, such as power on, power off, volume adjustment,), determine positions of the one or more microphones and the one or more cameras of the at least one peripheral device (see Vaughan Paragraph [0244], process may publish sensor data from the particular device to interface 804, such as sensor data from cameras, microphones, proximity sensors, wireless connections, and other sensor data relating to device states, positions, orientations, and other relevant context information, Paragraph [0245], sensor data received from various devices may indicate relative positions of devices and corresponding objects in an environment, such as device locations including a set-top box and corresponding television display, a home speaker), determine a position of the speaker connected with the audiovisual console (see Vaughan Paragraph [0255], As smartphone 912 comes into wireless communication range with home speaker 906, position information associated with smartphone 912 continues to be recorded as sensor data), regulate post-processing and a volume of the speaker (see Vaughan Paragraph [0265], The user may also initiate various functions which involve state changes at the device, such as power on, power off, volume adjustment), and regulate brightness and chrominance of a picture displayed on the display (see Vaughan Paragraph [0071], Graphics module 232 includes various known software components for rendering and displaying graphics on touch screen 212 or other display, including components for changing the visual impact (e.g., brightness, transparency, saturation, contrast, or other visual property) of graphics that are displayed).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of a system in which an audiovisual console connects to one or more peripheral devices through a communication interface, detects components of the peripheral device, obtains permission to access them, receives audio and video data, and generates audiovisual output to present as video on a display and audio to play on the speakers (as taught in Fitzpatrick), with a console that executes a control program that regulates audio and video by determining the positions of one or more microphones and one or more cameras, and a speaker, while adjusting audio processing, speaker volume, and display properties (as taught in Vaughan), the motivation being to provide a system that enables improved user control by integrating sensor data from devices into actionable information that allows the system to automatically optimize audiovisual settings and enhance user interactions within the environment (see Vaughan Paragraphs [0003] and [0027]).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Refer to PTO-892, Notice of References Cited for a listing of analogous art.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CARISSA A JONES whose telephone number is (703)756-1677. The examiner can normally be reached Telework M-F 6:30 AM - 4:00 PM CT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached at 5712727503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CARISSA A JONES/Examiner, Art Unit 2691
/DUC NGUYEN/Supervisory Patent Examiner, Art Unit 2691