DETAILED ACTION
Claim Rejections - 35 USC § 103
1. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
2. Claims 1-10, 12-24 and 26-32 are rejected under 35 U.S.C. 103 as being unpatentable over Angell et al, U.S. Patent Application Publication No. 2012/0063270 (hereinafter Angell) combined with Raft et al, U.S. Patent Application Publication No. 2022/0322001 (hereinafter Raft) in further view of Volgyesi et al, U.S. Patent Application Publication No. 2017/0328983 (hereinafter Volgyesi).
Regarding claim 1, Angel discloses a system (from abstract, see An event detection and localization system is provided that employs a plurality of smartphones) for locating an audio event comprising:
a wireless communication interface configured to communicate with audio device systems (from paragraph 0021, see The smartphones 110 also communicate in a known manner with one or more cellular base stations 140); and
a central processing unit coupled with the wireless communication interface (from paragraph 0021, see illustrates a network environment 100 where a plurality of smartphones 110-1 through 110-N interact with a server 120);
wherein the wireless communication interface is configured to receive audio device system signals and position data from the audio device systems, wherein the audio device system signals comprise data indicating local audio events respectively associated with the audio device systems (from paragraph 0032, see a notification (e.g., with the arrival time and arrival location) can be sent during step 380 to other smartphone nodes in the local environment and/or to a central server, if present); and
wherein the central processing unit is configured to:
determine whether an audio event has occurred based on the data indicating the local audio events respectively associated with the audio device systems (from paragraph 0033, see A test is performed during step 390, to determine if multiple notifications of the arrival time and arrival location of a potential gunshot are available from multiple smartphones. If it is determined during step 390 that multiple notifications are not available, then program control waits in step 390 until multiple notifications are available),
determine a position of the audio event based on the audio device system signals and the position data (from paragraph 0033, see the receiver calculates an implied origination of the detected gunshot during step 395), and
output the position of the audio event (from paragraph 0035, see The solution is then optionally presented to the users during step 398 on their respective smartphone display. If the smartphone has a compass, it can even provide a pointing approach to show the holder where the solution expects the origination point to be as well as providing a range).
Further regarding claim 1, Angell does not teach the audio device systems configured to be worn by respective users. All the same, Raft discloses the audio systems configured to be worn by respective users (from Figure 1, see 1). Therefore, it would have been obvious to one of ordinary skill in the art to modify Angell wherein the audio device systems configured to be worn by respective users as taught by Raft. This modification would have improved the system’s convenience by allowing the user to hold a weapon as suggested by Raft.
Again on the issue of claim 1, the combination of Angell and Raft does not clearly teach the central processing unit is configured to obtain a sound label, and determine whether the audio event has occurred based on the sound label. All the same, Volgyesi discloses the central processing unit is configured to obtain a sound label, and determine whether the audio event has occurred based on the sound label (from paragraph 0070, see Each sensor node then estimates ToA and AoA 728, 730, 732 for the received acoustic signal. Each sensor node also applies CNN analysis 734, 736, 738 to classify the transient event. The sensor nodes provide computed information to either a dedicated server 746 or a cloud computing infrastructure 742 for source localization computation. Results of the source localization computation may be presented through display 744). Therefore, it would have been obvious to one of ordinary skill in the art to further modify the combination of Angell and Raft wherein the central processing unit is configured to obtain a sound label, and determine whether the audio event has occurred based on the sound label as taught by Volgyesi. This modification would have improved the system’s reliability by reducing false positives as suggested by Volgyesi (see paragraph 0004).
Regarding claim 2, the combination of Angell and Raft discloses the audio device systems comprise hearing protection devices (from abstract of Raft see A hearing protection apparatus, a related method, and a server device is disclosed. The hearing protection apparatus comprises a set of microphones comprising a first microphone and a second microphone; a first positioning device for provision of first position data of the hearing protection apparatus; and a communication device connectable to the first microphone and the second microphone and the first positioning device, the communication device comprising a processing unit and an interface; wherein the processing unit is configured to: obtain a first audio input signal from the first microphone and a second audio input signal from the second microphone; determine first audio data based on the first audio input signal and the second audio input signal; and transmit the first audio data and the first position data to a radio unit).
Regarding claim 3, Angell discloses the audio device systems form a mesh network, and the central processing unit is communicatively coupled with the mesh network via the wireless communication interface (from paragraph 0032, see Wi-Fi or other wireless technology). In addition, the notification can be a multicast (one-to-many); a unicast (one-to-one) or a mesh approach (one smartphone notifying other smartphones, which, in turn, notify other smartphones). Various protocols can be used, such as UDP or TCP, depending on the network environment and whether there is a desire to confirm receipt to the various receivers).
Regarding claim 4, the combination of Angell and Raft discloses wherein one of the audio device system signals comprises microphone signals from one of the audio device systems (from paragraph 0088 of Raft, see such that the first audio input signal 29 and the second audio input signal are obtained from the hearing protection device 2).
Regarding claim 5, the combination of references as modified by Volgyesi discloses the sound label is associated with an input audio signal from one of the audio device systems (from paragraph 0021 of Volgyesi, see When an event is detected, a classification algorithm of an embodiment determines whether the event falls into one of the target classes or not).
Regarding claim 6, Angell discloses the system of claim 1, further comprising one of the audio device systems (from Figure 1, see 110-N).
Regarding claim 7, the combination of references as modified by Volgyesi discloses wherein the one of the audio device systems is configured to: determine a transient in an input audio signal, and determine one of the audio device system signals based on the transient (from abstract of Volgyesi, see A system is described that comprises a plurality of sensor nodes and at least one remote server, wherein each sensor node of the plurality of sensor nodes and the at least one remote server are communicatively coupled, wherein the plurality of sensor nodes receive at least one acoustic signal, process the at least one acoustic signal to detect one or more transient events, classify the one or more transient events as an event type, and determine geometry information and timing information of the one or more transient events. The system comprises at least one of the plurality of sensor nodes and the at least one remote server identifying the source of the one or more transient events).
Regarding claim 8, Angell discloses wherein the one of the audio device systems is configured to: determine a direction of arrival of an input audio signal, and determine one of the audio device system signals based on the direction of arrival (from paragraph 0033, see A test is performed during step 390, to determine if multiple notifications of the arrival time and arrival location of a potential gunshot are available from multiple smartphones. If it is determined during step 390 that multiple notifications are not available, then program control waits in step 390 until multiple notifications are available. If, however, it is determined during step 390 that multiple notifications are available, then the receiver calculates an implied origination of the detected gunshot during step 395. As more data is available, the solution can be refined or improved, as would be apparent to a person of ordinary skill in the art).
Regarding claim 9, the combination of Angell and Raft discloses wherein the one of the audio device systems is configured to: determine one or more timestamps for an audio input signal, and determine one of the audio device system signals based on the one or more timestamps (from paragraph 0040 of Raft, see The first audio data may comprise one or more time stamps such as one or more first time stamps related or associated with the first audio input signal and/or one or more second time stamps related or associated with the second audio input signal).
Regarding claim 10, the combination of Angell and Raft discloses wherein the one of the audio device systems comprises a microphone configured to provide an input audio signal, and wherein one of the audio device system signals from the one of the audio device systems comprises the input audio signal (from paragraph 0088 of Raft, see The hearing protection device interface 40 may allow the connection of the communication device 4 to the hearing protection device 2 of the hearing protection apparatus 13 of a mission-performing user 1, e.g. such that the first audio input signal 29 and the second audio input signal are obtained from the hearing protection device 2).
Regarding claim 12, the combination of references as modified by Volgyesi discloses wherein the central processing unit is configured to: determine a transient in an input audio signal from one of the audio device systems, and determine whether the audio event has occurred based on the transient (from abstract of Volgyesi, see A system is described that comprises a plurality of sensor nodes and at least one remote server, wherein each sensor node of the plurality of sensor nodes and the at least one remote server are communicatively coupled, wherein the plurality of sensor nodes receive at least one acoustic signal, process the at least one acoustic signal to detect one or more transient events, classify the one or more transient events as an event type, and determine geometry information and timing information of the one or more transient events. The system comprises at least one of the plurality of sensor nodes and the at least one remote server identifying the source of the one or more transient events).
Regarding claim 13, Angell discloses the system of claim 1, wherein the central processing unit is configured to: determine directions of arrival associated with the audio event, and determine the position of the audio event based on the directions of arrival (from paragraph 0033, see A test is performed during step 390, to determine if multiple notifications of the arrival time and arrival location of a potential gunshot are available from multiple smartphones. If it is determined during step 390 that multiple notifications are not available, then program control waits in step 390 until multiple notifications are available. If, however, it is determined during step 390 that multiple notifications are available, then the receiver calculates an implied origination of the detected gunshot during step 395. As more data is available, the solution can be refined or improved, as would be apparent to a person of ordinary skill in the art).
Regarding claim 14, Angell disclose the system of claim 1, wherein the central processing unit is configured to: determine times of arrival associated with the audio event, and determine the position of the audio event based on the times of arrival (from paragraph 0033, see A test is performed during step 390, to determine if multiple notifications of the arrival time and arrival location of a potential gunshot are available from multiple smartphones. If it is determined during step 390 that multiple notifications are not available, then program control waits in step 390 until multiple notifications are available. If, however, it is determined during step 390 that multiple notifications are available, then the receiver calculates an implied origination of the detected gunshot during step 395. As more data is available, the solution can be refined or improved, as would be apparent to a person of ordinary skill in the art).
Regarding claim 15, the combination of Angell and Raft discloses the central processing unit is configured to synchronize input audio signals from the respective audio device systems (from paragraph 0040 of Raft, see The first audio data may comprise one or more time stamps such as one or more first time stamps related or associated with the first audio input signal and/or one or more second time stamps related or associated with the second audio input signal).
Regarding claim 16, the combination of Angell and Raft discloses the central processing unit is deployable for a mission (from Figure 1 of Raft, see 1).
Regarding claim 17, the combination of Angell and Raft discloses the mission comprises a combat mission, and the central processing unit is deployable for the combat mission (from Figure 1 of Raft, see 1).
Regarding claim 18, the combination of Angell and Raft discloses the audio device system signals comprise microphone signals respectively from the audio device systems (from paragraph 0088 of Raft, see The hearing protection device interface 40 may allow the connection of the communication device 4 to the hearing protection device 2 of the hearing protection apparatus 13 of a mission-performing user 1, e.g. such that the first audio input signal 29 and the second audio input signal are obtained from the hearing protection device 2).
Regarding claim 19, Angell discloses wherein the audio device system signals comprise information regarding input audio signals respectively from the audio device systems (from paragraph 0033, see A test is performed during step 390, to determine if multiple notifications of the arrival time and arrival location of a potential gunshot are available from multiple smartphones. If it is determined during step 390 that multiple notifications are not available, then program control waits in step 390 until multiple notifications are available. If, however, it is determined during step 390 that multiple notifications are available, then the receiver calculates an implied origination of the detected gunshot during step 395. As more data is available, the solution can be refined or improved, as would be apparent to a person of ordinary skill in the art).
Regarding claim 20, Angell discloses an audio device system comprising:
one or more microphones configured to detect ambient sounds and to provide an input audio signal (from paragraph 0016, see Currently available smartphones typically incorporate a microphone for communications. If this microphone is maintained in a listening state, it can look for appropriate wave forms indicating a gunshot or another event);
a memory (from paragraph 0006, see Each smartphone comprises a memory for storing an event detection process);
a wireless communication interface (from paragraph 0032, see Bluetooth, cellular, Wi-Fi or other wireless technology);
a GPS module configured to determine first position data of the audio device system (from paragraph 0016, see a Global Positioning System (GPS) that allows a mobile device to determine a location for the particular mobile device); and
a processing unit (from paragraph 0042, see an integrated circuit, a digital signal processor, a microprocessor, and a micro-controller) configured to communicatively connect with another audio device system (from abstract, see another smartphone and a server);
wherein the processing unit is configured to:
receive an audio device system signal and position data from the other audio device system, wherein the audio device system signal comprise data indicating a local audio event associated with the other audio device system (from paragraph 0024, The second figure is for the isolated approach where the only available elements are the smartphones themselves. In this implementation, the phones all pass their available information to all other available smartphones and the smartphones serve as the environment to perform appropriate calculations and then exchange their solutions),
determine whether an audio event has occurred based on the data indicating the local audio event associated with the other audio device system, and based on the input audio signal (from paragraph 0033, see A test is performed during step 390, to determine if multiple notifications of the arrival time and arrival location of a potential gunshot are available from multiple smartphones. If it is determined during step 390 that multiple notifications are not available, then program control waits in step 390 until multiple notifications are available),
determine a position of the audio event (from paragraph 0033, see the receiver calculates an implied origination of the detected gunshot during step 395), and
output the position of the audio event (from paragraph 0035, see The solution is then optionally presented to the users during step 398 on their respective smartphone display. If the smartphone has a compass, it can even provide a pointing approach to show the holder where the solution expects the origination point to be as well as providing a range).
Further regarding claim 20, Angell does not clearly teach a speaker configured to provide an output audio signal. All the same, Raft discloses a speaker configured to provide an output audio signal (from Figure 2, see 44A and 46A). Therefore, it would have been obvious to one of ordinary skill in the art to further modify Angell with a speaker configured to provide an output audio signal as taught by Raft. This modification would have improved the system’s flexibility by allowing for sound alerts as suggested by Raft.
Again on the issue of claim 20, the combination of Angell and Raft does not clearly teach the central processing unit is configured to obtain a sound label, and determine whether the audio event has occurred based on the sound label. All the same, Volgyesi discloses the central processing unit is configured to obtain a sound label, and determine whether the audio event has occurred based on the sound label (from paragraph 0070, see Each sensor node then estimates ToA and AoA 728, 730, 732 for the received acoustic signal. Each sensor node also applies CNN analysis 734, 736, 738 to classify the transient event. The sensor nodes provide computed information to either a dedicated server 746 or a cloud computing infrastructure 742 for source localization computation. Results of the source localization computation may be presented through display 744). Therefore, it would have been obvious to one of ordinary skill in the art to further modify the combination of Angell and Raft wherein the central processing unit is configured to obtain a sound label, and determine whether the audio event has occurred based on the sound label as taught by Volgyesi. This modification would have improved the system’s reliability by reducing false positives as suggested by Volgyesi (see paragraph 0004).
Regarding claim 21, Angell discloses the audio device system of claim 20, wherein the audio device system and the other audio device system are parts of a wireless network (from paragraph 0032, see Bluetooth, cellular, Wi-Fi or other wireless technology). In addition, the notification can be a multicast (one-to-many); a unicast (one-to-one) or a mesh approach).
Regarding claim 22, Angell discloses the audio device system of claim 20, wherein the audio device system and the other audio device system are parts of a mesh network (from paragraph 0032, see Bluetooth, cellular, Wi-Fi or other wireless technology). In addition, the notification can be a multicast (one-to-many); a unicast (one-to-one) or a mesh approach).
Regarding claim 23, the combination of Angell and Raft discloses the audio device system comprises a hearing protection device (from abstract of Raft see A hearing protection apparatus, a related method, and a server device is disclosed. The hearing protection apparatus comprises a set of microphones comprising a first microphone and a second microphone; a first positioning device for provision of first position data of the hearing protection apparatus; and a communication device connectable to the first microphone and the second microphone and the first positioning device, the communication device comprising a processing unit and an interface; wherein the processing unit is configured to: obtain a first audio input signal from the first microphone and a second audio input signal from the second microphone; determine first audio data based on the first audio input signal and the second audio input signal; and transmit the first audio data and the first position data to a radio unit).
Regarding claim 24, the combination of Angell and Raft discloses the one or more microphones are hear-through and/or feedforward microphones (from abstract of Raft, see A hearing protection apparatus, a related method, and a server device is disclosed. The hearing protection apparatus comprises a set of microphones comprising a first microphone and a second microphone; a first positioning device for provision of first position data of the hearing protection apparatus; and a communication device connectable to the first microphone and the second microphone and the first positioning device, the communication device comprising a processing unit and an interface; wherein the processing unit is configured to: obtain a first audio input signal from the first microphone and a second audio input signal from the second microphone; determine first audio data based on the first audio input signal and the second audio input signal; and transmit the first audio data and the first position data to a radio unit).
Regarding claim 26, the combination of references as modified by Volgyesi discloses wherein the processing unit is configured to determine a transient in the input audio signal (from abstract, see A system is described that comprises a plurality of sensor nodes and at least one remote server, wherein each sensor node of the plurality of sensor nodes and the at least one remote server are communicatively coupled, wherein the plurality of sensor nodes receive at least one acoustic signal, process the at least one acoustic signal to detect one or more transient events, classify the one or more transient events as an event type, and determine geometry information and timing information of the one or more transient events. The system comprises at least one of the plurality of sensor nodes and the at least one remote server identifying the source of the one or more transient events).
Regarding claim 27, Angell discloses the audio device system of claim 20, wherein the processing unit is configured to determine a direction of arrival associated with the input audio signal (from paragraph 0033 of Raft, see A test is performed during step 390, to determine if multiple notifications of the arrival time and arrival location of a potential gunshot are available from multiple smartphones).
Regarding claim 28, the combination of Angell and Raft discloses wherein the processing unit is configured to determine one or more timestamps for the input audio signal (from paragraph 0040 of Raft, see The first audio data may comprise one or more time stamps such as one or more first time stamps related or associated with the first audio input signal and/or one or more second time stamps related or associated with the second audio input signal).
Regarding claim 29, the combination of references as modified by Volgyesi discloses the central processing unit is configured to obtain the sound label from one of the audio device systems (from paragraph 0070 of Volgyesi, see The sensor nodes provide computed information to either a dedicated server 746 or a cloud computing infrastructure 742 for source localization computation).
Regarding claim 30, the combination of references as modified by Volgyesi discloses the sound label is associated with an input audio signal and is different from the input audio signal (from paragraph 0021 of Volgyesi, see When an event is detected, a classification algorithm of an embodiment determines whether the event falls into one of the target classes or not).
Regarding claim 31, the combination of references as modified by Volgyesi discloses the processing unit is configured to provision another sound label based on the input audio signal (from paragraph 0031 of Volgyesi, see The CNN may be trained with a large set of labeled examples, and the result of this supervised learning process - the inferred weights and biases - is embedded in each sensor node. This approach, most notably the separation of software code and trained data, also enables a straightforward and simplistic mechanism for improving and updating the classifier in already deployed equipment, since no software updates are required, only new data models. The current CNN architecture depends under one embodiment on 2-4 million weights, thus requiring sizeable and diverse training datasets).
Regarding claim 32, the combination of references as modified by Volgyesi discloses the audio system is configured to transmit the other sound label to the other audio device system (from paragraph 0031 of Volgyesi, see The CNN may be trained with a large set of labeled examples, and the result of this supervised learning process - the inferred weights and biases - is embedded in each sensor node. This approach, most notably the separation of software code and trained data, also enables a straightforward and simplistic mechanism for improving and updating the classifier in already deployed equipment, since no software updates are required, only new data models. The current CNN architecture depends under one embodiment on 2-4 million weights, thus requiring sizeable and diverse training datasets).
Response to Arguments
3. Applicant’s arguments have been considered but are deemed to be moot in view of the new grounds of rejection.
Conclusion
4. Applicant’s amendment necessitated the new ground(s) of rejection presented in this Office action. THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
5. Any inquiry concerning this communication or earlier communications from the examiner should be directed to OLISA ANWAH whose telephone number is 571-272-7533. The examiner can normally be reached Monday to Friday from 8.30 AM to 6 PM.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Carolyn Edwards can be reached on 571-270-7136. The fax phone numbers for the organization where this application or proceeding is assigned are 571-273-8300 for regular communications and 571-273-8300 for After Final communications.
Any inquiry of a general nature or relating to the status of this application or proceeding should be directed to the receptionist whose telephone number is 571-272-2600.
Olisa Anwah
Patent Examiner
August 1, 2026
/OLISA ANWAH/Primary Examiner, Art Unit 2692