DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Examiner’s Comments
Prior art to Lee et al (US 20160104491 A1) teaches that audio video devices can comprise audio source separation from a single mixed signal where the objects are separated into object types as shown in fig. 7.
Drawings
The drawings are objected to under 37 CFR 1.83(a). The drawings must show every feature of the invention specified in the claims. Therefore, the one or more spatialization filters (alternative to the HRTF filters) of claim 28; the smartphone, mobile phone, tablet, laptop, personal computer, set-top box, smart speaker, soundbar, handheld device, vehicle system, and headphones headset in ear soundbar speaker array or speakers of claim 29, must be shown or the feature(s) canceled from the claim(s). No new matter should be entered.
Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claim 27-61 rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-20 of U.S. Patent No. 11611840. Although the claims at issue are not identical, they are not patentably distinct from each other because the application claim 27 is a broader version of the of the patent claim 1 with obvious variations.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
The following claims 27-40 is/are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Horrocks et al (US 20200137494 A1).
As per claim 27, a sound generation system, comprising:
two or more devices (the microphones on the devices in fig. 1),
wherein at least one device functions as a sound source device, and one or more devices function as sound receiving devices (the devices in communication with each other),
each device being connected via wired or wireless communication (fig. 1), and
each device comprising of one or more processors (required to implement the disclosed terminals) and
where at least one of the sound source device or the one or more sound receiving devices comprise one or more sensors (the microphones and means to record users voices),
with the sound receiving devices further comprising one or more sound generation components (the means to playback the disclosed audio streams including the 3d mixed audio in fig. 1);
wherein the sound source device is configured to transmit an audio stream (each device transmits audio streams), and
the sound receiving devices are configured to receive the audio stream from the sound source device (each device receives the audio streams);
wherein a relative location, position, distance, angle or orientation between the sound source device and the sound receiving devices is determined based at least in part on data provided by the one or more sensors (the microphones record the voices of the users used in the communications to make the audio streams which are correlated with a location from another sensor per para 28 );
wherein the system is configured to process the audio stream based on said relative locations position, distance, angle or orientation, such that a listener using the sound receiving devices perceives the spatial sound (per para 28),
including cues indicating the sound as originating from the direction of the sound source device (per the 3d mixed audio in view of para 28).
As per claim 28, the sound generation system of claim 27, wherein the one or more devices are configured to: retrieve head-related transfer functions (HRTF) or one or more spatialization filters based on the determined relative location position, distance, angle or orientation between the sound source device and the sound receiving devices (the HRTF rejection must be implemented/retrieved per the location in order to function per, para 19);
process the audio stream using the retrieved HRTF or one or more spatialization filters (applying the determined/retrieved HRTF to the audio streams); and generate two or more audio channels to drive the sound generation components in the sound receiving devices (output of the 3d mix in fig. 1).
As per claim 29, the sound generation system of claim 27, wherein said sound source device is a smartphone, mobile phone, tablet, laptop, personal computer, set-top box, smart speaker, soundbar, handheld device, vehicle system, TV, or computer, and said one or more receiving devices are earphones, earbuds, or speakers (as shown in fig. 1).
As per claim 30, a method of generating sound using a system comprising two or more devices, the method comprising: designating at least one device as a sound source device and one or more devices as sound receiving devices (each of the devices in fig. 1 sends and receives data per protocols designating them as such) , each device being connected via wired or wireless communication (fig. 1), and each device comprising one or more processors and where at least one of the sound source device or the one or more sound receiving devices comprise one or more sensors (required to perform the functions per the claim 27 rejection), with the sound receiving devices further comprising one or more sound generation components (the speakers and microphones used for communication per fig. 1); transmitting an audio stream from the sound source device to the one or more sound receiving devices (each device sends and receives when in communication); receiving said audio stream at the one or more sound receiving devices (as shown in fig. 1); determining the relative location position distance angle or orientation between said sound source device and said sound receiving devices based on data provided by said sensors (the information used to create the 3d mix in fig. 1 is based on the audio streams captured by sensors/microphones at each devices used to capture speech to create the audio streams); and processing the audio stream based on the determined relative location, position distance angle or orientation such that a listener using the sound receiving devices perceives the spatial sound (per the claim 27 rejection), including cues indicating the sound as originating from the direction of said sound source device (per claim 27 rejection).
As per claim 31, the method of claim 30, further comprising: retrieving head-related transfer functions (HRTF) or one or more other spatialization filters based on the determined relative location, position distance angle or orientation between said sound source device and said sound receiving devices (the information used to create the 3d mix in fig. 1); processing the audio stream using the retrieved HRTF or the spatialization filters (per claim 27 rejection as applied to the audio streams); and generating two or more audio channels to drive the sound generation components in the sound receiving devices (the 3d audio mix per fig. 1).
As per claim 32, the method of claim 30, further comprising: using a sound source device selected from the group consisting of a smartphone, mobile phone, laptop, TV, or computer (the devices in fig 1 each comprise a computer), and transmitting the sound to one or more receiving devices selected from the group consisting of earphones, earbuds, or speakers (the cans 35 and 37 are earphones/earbuds and speakers).
As per claim 33, a sound generation system, comprising: one or more processing devices configured to:
receive a single mixed sound/audio stream comprising audio of a plurality of sound types (as received by each of the devices of fig. 1 when in communication);
receive a specification, pre-configuration, default configuration, or user configuration of one or more desired output sound types and corresponding volumes for each sound type (any of the parameters used by the processor to process the audio stream in the devices per the claim 27 rejection);
separate the single mixed sound/audio stream into one or more separated sound tracks corresponding to the plurality of sound types (corresponding to because they derive from the same prior audio processing as cited above) based on the received specification or configuration (the separation of tracks required to make the 3d mix in fig. 1 based on the step in para 19: processes incoming audio sources to introduce effects which induce the listener, in this example call participant 30, to perceive positional location for each connected audio stream ) (noting that each audio stream must be separated in order to apply the respective positional location for each source); and
adjust the volume of one or more separated sound tracks (the volumes of the sound sources must be set in order to present them to a particular position per the 3d mix) and
combine or mix the sound tracks based on the specification or configuration to generate one or more audio channels to drive one or more sound generation devices (the 3d mix in fig. 1).
As per claim 34, the sound generation system of claim 33, wherein each output audio channel is configured to output a single sound type selected from the group consisting of solo, vocals only, vocals combined with instrumental sounds, instrumental sounds only, voice only, music only, or a weighted combination of one or more sound types (the audio channel output the vocals of each participant per fig. 1).
As per claim 35, the sound generation system of claim 33, wherein the volume of each audio channel and the weighted combination of sound types are controlled by either user input or the received specification or configuration, allowing the user to individually control or adjust the volume of different sound types (the configuration parameters used to perform the 3d mixed sound per fig. 1) (further, the claim is not mapped as weighted combination is recited in the alternative in the parent claim).
As per claim 36, the sound generation system of claim 33, wherein the separation of the sound stream into multiple sound tracks and the generation of the audio channels are performed by one or more systems selected from the group consisting of machine learning systems, artificial neural network systems, digital signal processing systems, or any combination thereof (the devices of fig. 1 each comprise dsps in order to perform the functions of the claim 27 rejection).
As per claim 37, A method for generating sound, the method comprising: receiving, by one or more processing devices, a single mixed audio stream comprising audio of a plurality of sound types (the devices in communication in fig. 1 per the claim 27 rejection corresponding to voice and silence/noise types);
receiving a specification, pre-configuration, default configuration, or user configuration of one or more desired output sound types and corresponding volumes for each sound type (the parameters required to makes to corresponding gain differences/volumes between the left and right output channels of the cans 35 and 37 to make the 3d mix of fig. 1);
separating, by the one or more processing devices, the sound stream/single mixed audio into one or more separated sound tracks, each one corresponding to one of the plurality of sound types based on the received specification or configuration (creating the different sources for the 3d mixed sound, where each source is a type of sound, ie. Voice versus silence versus noise); and
adjusting the volume of one or more separated sound tracks and combining or mixing the separated sound tracks based on the specification or configuration to generate one or more audio channels to drive one or more sound generation devices (using the HRTF and stereo processing to produce 3d signals for the cans 35 and 37 in fig. 1).
As per claim 38, the method of claim 37, further comprising: configuring each output audio channel to output a single sound type selected from the group consisting of solo, vocals only, vocals combined with instrumental sounds, instrumental sounds only, voice only, music only, or a weighted combination of one or more sound types (the devices output voices only/call participants for the 3d mixed sound for para 34 ).
As per claim 39, the method of claim 37, further comprising: adjusting the volume of each audio channel and the weighted combination of sound types by either user input or the received specification or configuration, allowing the user to individually control or adjust the volume of different sound types (per the claim 35 rejection) (further, the claim is not mapped as weighted combination is recited in the alternative in the parent claim).
As per claim 40, the method of claim 37, wherein: the separation of the sound stream into multiple sound tracks and the generation of the audio channels are performed by one or more systems selected from the group consisting of machine learning systems, artificial neural network systems, digital signal processing systems, or any combination thereof (the functions in the claim 27 rejection require dsps in each of the terminal devices).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The following claims 41-45 is/are rejected under 35 U.S.C. 103 as being unpatentable over Horrocks et al (US 20200137494 A1), and further in view of Lee et al (US 20160104491 A1).
As per claim 41, Horrocks discloses a sound generation system, comprising:
one or more processing devices (the terminals in fig. 1), one or more display devices (the displays to display the gui 3 on the devices in fig. 1), and one or more sound generation devices (the cans 35,37 in fig. 1); wherein the system is configured to: receive a video stream comprising visual and audio components, the audio components including audio signals and the visual components including video images (part of the video conference per para 49);
recognize a sound field from the video stream and generate a configuration of the sound field (the parameters with the audio streams used to create the 3d mixed sound in fig. 1), the configuration including one or more sound sources, sound source locations, and a listener's position (per the claim 27 rejection, the sound source positions are all relative to a central user position in the same 3d coordinate system);
However, Horrocks does not specify:
wherein at least one of the sound source locations corresponds to a location of the corresponding sound source displayed on one or more of the display devices;
separate the audio component into one or more sound tracks, each sound track corresponding to one or more sound sources;
generate one or more audio channels by combining or mixing the separated sound tracks according to the configuration of the sound field; and
utilize the generated audio channels to drive the sound generation devices, thereby generating a sound field such that one or more users or listeners perceive sound of each sound source as originating from the corresponding location on the display.
Lee teaches that audio video devices can comprise audio source separation from a single mixed signal where the objects are separated into object types as shown in fig. 7. Lee further teaches to implement speakers on a screen in order to enable sound matching each object position in the screen per para 167. It would have been obvious to one skilled in the art at the time of filing to implement the speakers per para 187 into the tv of Horrocks for the purpose of allowing audio that tracks the video on the television.
As such the combined system comprises:
wherein at least one of the sound source locations corresponds to a location of the corresponding sound source displayed on one or more of the display devices;(Lee para 167)
separate the audio component into one or more sound tracks, each sound track corresponding to one or more sound sources (Lee Fig. 7);
generate one or more audio channels by combining or mixing the separated sound tracks according to the configuration of the sound field (Horrocks per the claim 27 rejection and/or Lee Fig. 7); and
utilize the generated audio channels to drive the sound generation devices, thereby generating a sound field such that one or more users or listeners perceive sound of each sound source as originating from the corresponding location on the display (The Gui of Horrocks and as per the claim 27 rejection, in view of the screen-audio object tracking taught by Lee).
As per claim 42, the sound generation system of claim 41, wherein the recognition of one or more sound sources and sound source locations in the video images, and the separation of the audio component into one or more sound tracks associated with the sound sources, are performed by one or more machine learning systems, one or more artificial neural network systems, one or more digital signal processing systems, or any combination thereof (the terminals each require a dsp to perform the functions cited above).
As per claim 43, the sound generation system of claim 41, wherein the separated audio tracks are processed using head-related transfer functions (HRTF) or one or more spatialization filters (per the claim 27 rejection), and the processed audio tracks are combined or mixed to generate the audio channels that drive the sound generation devices, such as earbuds or earphones, thereby enabling listeners to perceive a sound field where the sound of each sound source is perceived as originating from the corresponding location on the display (per fig. 1 and fig. 3 and per the claim 27 rejection).
As per claim 44, the sound generation system of claim 41, wherein when a displayed or configured sound source location is changed, the corresponding sound source location in the sound field changes dynamically (fig. 3 and para 32, and as per the object tracking taught by Lee above).
As per claim 45, the sound generation system of claim 33, wherein the single mixed audio stream is an audio component of audio content, video content, or broadcast content (per the 27 and 41 rejections).
The following claims 46,47,49,50,51,53-57 is/are rejected under 35 U.S.C. 103 as being unpatentable over Horrocks et al (US 20200137494 A1) in view of Lee et al (US 20160104491 A1), and further in view of Oh et al (US 20140222440 A1).
As per claim 46, the sound generation system of claim 33, Horrocks and Lee disclose the sound generation system of claim 33, but do not specify:
wherein the plurality of sound types comprises a voice sound type and a background sound type, and wherein the one or more processing devices are configured to independently adjust a volume of the voice sound type relative to a volume of the background sound type or to adjust a ratio of the volume of the voice sound type to the volume of the background sound type.
Oh teaches that audio decoders and classify and render objects and background objects independently, each with adjustable volume per the abstract: the mix information being usable to control gain or panning of the independent object or the background object. It would have been obvious to one skilled in the art at the time of filing that the objects of Horrocks and Lee could comprise background objects and objects with adjustable levels as taught by Oh for the purpose of allowing for processing an audio signal that can control the gain and panning of an object without limitation (per para 8).
As per claim 47, the sound generation system of claim 46, wherein the one or more processing devices are configured to increase the volume of the voice sound type and decrease the volume of the background sound type to enhance intelligibility of the voice sound type, wherein enhancing intelligibility of the voice sound type is also appliable to noise reduction, noise cancellation, denoising, or any combination thereof (per the adjustable independent background/noise object and the object in the claim 46 rejection).
As per claim 49, the method of claim 37, wherein the single mixed audio stream is an audio component of audio content, video content, or broadcast content (per the above rejections).
As per claim 50, the method of claim 37, wherein the plurality of sound types comprises a voice sound type and a background sound type, the method further comprising independently adjusting a volume of the voice sound type relative to a volume of the background sound type or to adjust a ratio of the volume of the voice sound type to the volume of the background sound type (per claim 47 rejection).
As per claim 51, the method of claim 50, further comprising increasing the volume of the voice sound type and decreasing the volume of the background sound type to enhance intelligibility of the voice sound type, wherein enhancing intelligibility of the voice sound type is also applicable to noise reduction, noise cancellation, denoising, or any combination thereof (per claim 48 rejection).
As per claim 53, A non-transitory machine-readable medium storing instructions that, when executed by one or more processing devices (the system disclosed in the above rejections requires processors, memory and software in order to be implemented) of a system comprising two or more devices, cause the one or more processing devices to: designate at least one device as a sound source device and
one or more devices as sound receiving devices, each device being connected via wired or wireless communication, wherein at least one of the sound source device or the sound receiving devices (terminals in fig. 1 of Horrocks)
comprises one or more sensors, and the sound receiving devices further comprise one or more sound generation components; cause the sound source device to transmit an audio stream to the one or more sound receiving devices; determine a relative location, position, distance, angle, or orientation between the sound source device and the sound receiving devices based at least in part on data provided by the one or more sensors; and process the audio stream based on the determined relative location, position, distance, angle, or orientation, so that a listener using the sound receiving devices perceives spatial sound including cues indicating the sound as originating from the direction of the sound source device (per claim 27 rejection).
As per claim 54, the non-transitory machine-readable medium of claim 53, wherein the
operations further comprise retrieving head-related transfer functions (HRTF) or one or more other
spatialization filters based on the determined relative location, position, distance, angle, or
orientation, processing the audio stream using the retrieved HRTF or one or more other
spatialization filters, and generating two or more audio channels to drive the sound generation
components (per the HRTF processing in view of the location/position processing cited in the above rejections).
As per claim 55, a non-transitory machine-readable medium storing instructions that, when
executed by one or more processing devices, cause the one or more processing devices to: receive
a single mixed audio stream comprising audio of a plurality of sound types; receive a specification,
pre-configuration, default configuration, or user configuration of one or more desired output sound
types and corresponding volumes for each sound type; separate the single mixed audio stream into
one or more separated sound tracks, each separated sound track corresponding to one of the
plurality of sound types, based on the received specification or configuration; and adjust a volume
of one or more of the separated sound tracks and combine or mix the separated sound tracks based
on the specification or configuration to generate one or more audio channels to drive one or more
sound generation devices
(per the claim 28 rejection above, noting the sound types and separated objects of Lee).
As per claim 56, the non-transitory machine-readable medium of claim 55, wherein the plurality
of sound types comprises a voice sound type and a background sound type, and wherein the
operations further comprise independently adjusting a volume of the voice sound type relative to
a volume of the background sound type.
( per claim 46,47,50,51 rejections above).
As per claim 57, the sound generation system of claim 27, wherein the one or more sensors comprise at least one of an accelerometer, a gyroscope, a position sensor (any of sensors per para 28-31. Notably: or user interface 95 may provide an indication of each person on the call; or a position for each interaction, audio stream or person), a motion sensor, a magnetometer, or a global positioning system sensor.
The following claims 58,59,48,52,60,61 is/are rejected under 35 U.S.C. 103 as being unpatentable over Horrocks et al (US 20200137494 A1) in view of Lee et al (US 20160104491 A1), and further in view of Oh et al (US 20140222440 A1) and further in view of Stein (US 20170366914 A1).
As per claim 58, Horrocks Lee and Oh disclose a sound generation system, comprising: one or more processing devices
configured to: represent a listener and one or more sound sources in a three-dimensional sound
field, a relative location between the listener and each sound source being represented by at least one of an azimuth angle, an attitude angle, or a range (the location position and HRTF filtering as cited above, noting that all parameters supporting the cited functions represent a range of values defined by the quantization levels of that particular processor);
retrieve, for each sound source, one or more head-related transfer functions (HRTF) or other spatialization filters based on the relative location (para 19: For example, a head-related transfer function (HRTF), or anatomical transfer function (ATF), which characterizes how an ear receives a sound from a point in space, may be used to synthesize a binaural sound that seems to come from a particular point in space where the HRTF is adapted for the various positions);
process audio of the sound source using the retrieved HRTF or other spatialization filters to generate one or more audio channels (para 25 per the 3d space);
However Horrocks in view of Lee in view of Oh do not disclose
To process audio responsive to movement of at least one of the listener;
the sound source that changes the relative location, dynamically retrieve one or more updated HRTF or other spatialization filters and update the one or more audio channels so that the sound source is perceived as remaining at its location in the three-dimensional sound field as the listener moves.
Stein teaches that audio devices can use sensors to determine listener position in order to adjust a rendered soundfield in order to maintain the object position while the user moves (para 114 to 116 via the maintained sound field with the tracked user position per para 150). It would have been obvious to one skilled in the art at the time of filing, to implement listener tracking to adapt the rendering for the purpose of processing such that the audio maintains an absolute alignment with the intended sound scene or visual components as cited in para 150).
As per claim 59, the sound generation system of claim 58, wherein at least one of the one or more sound sources is associated with a sound source device, the movement of the listener is
determined based at least in part on data provided by one or more sensors, and the audio is processed so that the sound source associated with the sound source device is perceived as originating from a direction of the sound source device (per the sensors and soundfield cited in the claim 58 rejection).
As per claim 48, Horrocks et al (US 20200137494 A1) in view of Lee et al (US 20160104491 A1), and further in view of Oh et al (US 20140222440 A1) disclose the sound generation system of claim 33, but do not specify further comprising a speech recognition unit configured to recognize a voice command, wherein at least one of the separation or the volume adjustment is controlled at least in part by the recognized voice command.
The examiner takes official notice it is well known in the art to use voice commands via microphone arrays; where it would be obvious to one skilled in the art at the time of filing to implement voice commands with microphone arrays for the purpose of improved user interfaces.
As per claim 52, the method of claim 37, further comprising recognizing a voice command with a speech recognition unit, and controlling at least one of the separating or the volume adjusting based at least in part on the recognized voice command (the GUI of Horrocks in view of the voice commands per the claim 48 rejection).
As per claim 60, the sound generation system of claim 48, wherein the voice command is
captured by a microphone array comprising a plurality of microphones (per claim 48 rejection).
As per claim 61, the method of claim 52, wherein the voice command is captured by a
microphone array comprising a plurality of microphones (per claim 48 rejection).
Response to Arguments
The submitted arguments have been considered but are moot in view of the new grounds of rejection.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDER KRZYSTAN whose telephone number is 571-272-7498, and whose email address is alexander.krzystan@uspto.gov
The examiner can usually be reached on m-f 7:30-4:00 est.
If attempts to reach the examiner by telephone or email are unsuccessful, the examiner’s supervisor, Carolyn Edwards can be reached on (571) 270-7136.
The fax phone numbers for the organization where this application or proceeding is assigned are 571-273-8300 for regular communications and 571-273-8300 for After Final communications.
/ALEXANDER KRZYSTAN/Primary Examiner, Art Unit 2653
Examiner Alexander Krzystan
September 1, 2026