DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 03/20/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over CN 116206624 A, published June 2nd, 2023, hereinafter “Zhao”, in view of US 20250330262 A1, with an earliest priority date of April 18th, 2024, hereinafter “Kim”.
Regarding claim 1, Zhao teaches a method of controlling sound using a sound quality index-based generative artificial intelligence (Al), the method comprising: inputting, by a controller, sound generated from a vehicle to a first encoder. See at least [0107], [n0052]-[n0054], and figure 3, wherein noise signals emitted by an engine of a vehicle are input into a parameter encoding layer (VAE) of an initial vehicle sound synthesis model. See at least [0062]-[0063], wherein the disclosed methods are performed by a processor.
and outputting sound through a first decoder based on the first latent vector. See at least [n0055]-[n0058] and figure 3, wherein the encoding from the first encoder is used by a decoding layer of the model to generate an output noise signal.
Zhao remains silent on the encoders being neural, and generating a first latent vector for the sound in the first neural encoder.
Kim teaches the encoders being neural. See at least [0052]-[0053] and figure 4, wherein the encoders and decoders are performed by neural networks.
generating a first latent vector for the sound in the first neural encoder. See at least [0110]-[0111] and figure 7, step S130, wherein a first encoding neural network performs an encoding operation on the input data to generate first latent data. See at least [0084], wherein the generated latent data comprises encoded code vectors. See at least [0003]-[0005] and [0053], wherein the input data is audio signal data.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to modify Zhao with Kim’s technique of neural encoders, and generating a first latent vector for the sound in the first neural encoder. It would have been obvious to modify because doing so enables improvements in data transmission, in particular during compression and restoration of audio data, as recognized by Kim (see at least [0003]-[0007]).
Regarding claim 2, Zhao and Kim in combination teach all of the limitations of claim 1 as discussed above, and Zhao additionally teaches wherein the controller is configured to simultaneously train the first neural encoder and the first neural decoder by comparing the output sound with target sound. See at least [n0050], wherein the sound synthesis model, including the encoding layer and the decoding layer, is trained by comparing the output synthesis results with calibration sound results with target loss functions.
Regarding claim 3, Zhao and Kim in combination teach all of the limitations of claim 2 as discussed above, and Zhao additionally teaches wherein the sound received at the first neural encoder is generated from an engine or a motor of the vehicle. See at least [0107], wherein the collected sound signals for input are generated from the engine through the exhaust system of the vehicle as the vehicle is driving.
Regarding claim 20, Zhao teaches An apparatus for controlling sound using a sound quality index-based generative artificial intelligence (Al), the apparatus comprising: a controller configured to input sound generated from a vehicle to a first encoder. See at least [0107], [n0052]-[n0054], and figure 3, wherein noise signals emitted by an engine of a vehicle are input into a parameter encoding layer of an initial vehicle sound synthesis model. See at least [0062]-[0063], wherein the disclosed methods are performed by a processor.
wherein the controller configured to output sound through a first decoder based on the first latent vector. See at least [n0055]-[n0058] and figure 3, wherein the encoding from the first encoder is used by a decoding layer of the model to generate an output noise signal.
Zhao remains silent on the encoders being neural, and generate a first latent vector for the sound in the first neural encoder.
Kim teaches the encoders being neural. See at least [0052]-[0053] and figure 4, wherein the encoders and decoders are performed by neural networks.
generate a first latent vector for the sound in the first neural encoder. See at least [0110]-[0111] and figure 7, step S130, wherein a first encoding neural network performs an encoding operation on the input data to generate first latent data. See at least [0084], wherein the generated latent data comprises encoded code vectors. See at least [0003]-[0005] and [0053], wherein the input data is audio signal data.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to modify Zhao with Kim’s technique of neural encoders, and generating a first latent vector for the sound in the first neural encoder. It would have been obvious to modify because doing so enables improvements in data transmission, in particular during compression and restoration of audio data, as recognized by Kim (see at least [0003]-[0007]).
Claims 4-19 are rejected under 35 U.S.C. 103 as being unpatentable over Zhao and Kim in combination as applied to claims above, and further in view of US 20200193960 A1, published June 18th, 2020, hereinafter “Jung”.
Regarding claim 4, Zhao and Kim in combination teach all of the limitations of claim 3 as discussed above, and Zhao additionally teaches inputting, by a controller, vehicle information to a second neural encoder. See at least [0093], [n0040]-[n0045], and figures 2-3, wherein vehicle parameters are input to a second encoding layer (paramVAE, separate from first encoding layer VAE).
and outputting sound through a second neural decoder based on the second latent vector, wherein the second neural decoder is configured to use the first neural decoder of which training is completed. See at least [n0045]-[n0048], wherein the output of the second encoder paramVAE is used by a decoding layer to output a sound signal. See at least [n0072], wherein the decoding layer of the trained vehicle sound synthesis model (second decoder, figure 2) is formed based on the decoding layer of the initial vehicle sound synthesis model (first decoder, figure 3). See at least [n0014] and [0040]-[0054], wherein the synthesis unit includes two decoding layers, one in the synthesis unit and one in a training subunit of the synthesis unit.
Zhao remains silent on a controller area network (CAN) signal and a vibration signal of the vehicle, the encoders being neural encoders, and generating a second latent vector for the CAN signal and the vibration signal processed at the second neural encoder.
Kim teaches the encoders being neural encoders, and generating a second latent vector for the signals processed at the second neural encoder. See at least [0110]-[0111] and figure 7, step S130, wherein multiple neural networks perform an encoding operations on the input data to generate latent data. See at least [0084], wherein the generated latent data comprises encoded code vectors. See at least [0003]-[0005] and [0052]-[0053], wherein the input data is audio signal data, and the encoders and decoders are performed by neural networks.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to modify Zhao with Kim’s technique of neural encoders, and generating latent vectors for input data processed at multiple neural encoders. It would have been obvious to modify because doing so enables improvements in data transmission, in particular during compression and restoration of audio data, as recognized by Kim (see at least [0003]-[0007]).
Jung teaches a controller area network (CAN) signal and a vibration signal of the vehicle. See at least [0081], [0096], and [0147], wherein vehicle driving information 411 is input from a controller area network of the vehicle, along with engine characteristic information 310. See at least [0118], wherein engine characteristic information 310 includes vibration signal information.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to further modify Zhao with Jung’s technique of collecting CAN signals and vibration signals of the vehicle. It would have been obvious to modify because doing so enables optimized vehicle engine tone using artificial intelligence, as recognized by Jung (see at least [0006]-[0008]).
Regarding claim 5, Zhao, Kim, and Jung in combination teach all of the limitations of claim 4 as discussed above, and Zhao remains silent on the specifics of wherein a loss function is based on a difference between the first latent vector for the sound generated of the vehicle and the second latent vector for the CAN signal and vibration signal of the vehicle. However, Zhao does teach a loss function based on a difference. See at least [n0064], wherein the first loss function measures the difference between the encoded data provided by the parameter encoding layer and the decoded data provided by the decoding layer.
Kim teaches wherein a loss function is based on a difference between the first latent vector for the sound generated of the vehicle and the second latent vector for the CAN signal and vibration signal of the vehicle. See at least [0165]-[0166], wherein the difference loss function is based on a difference between the latent data of the ith encoder and the jth encoder.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to modify Zhao with Kim’s loss function based on a difference between multiple latent vectors from multiple encoders. It would have been obvious to modify because doing so enables improvements in data transmission, in particular during compression and restoration of audio data, as recognized by Kim (see at least [0003]-[0007]).
Regarding claim 6, Zhao, Kim, and Jung in combination teach all of the limitations of claim 5 as discussed above, and Zhao additionally teaches wherein the controller is configured to train the second neural encoder by comparing the sound output from the second neural decoder with target sound. See at least [n0050] and [n0072], wherein the second encoder (paramVAE) is trained by comparing synthesized output sound with a target calibration sound signal.
Regarding claim 7, Zhao, Kim, and Jung in combination teach all of the limitations of claim 6 as discussed above, and Zhao remains silent on further comprising training a first neural network configured to receive the output sound to thereby output a sound quality index (SQI) and compare the sound quality index with a target sound quality index.
Kim teaches a first neural network. See at least [0055]-[0056], first neural network NN1.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to modify Zhao with Kim’s neural network. It would have been obvious to modify because doing so enables improvements in data transmission, in particular during compression and restoration of audio data, as recognized by Kim (see at least [0003]-[0007]).
Jung teaches further comprising training a first artificial intelligence configured to receive the output sound to thereby output a sound quality index (SQI) and compare the sound quality index with a target sound quality index. See at least [0070] and [0162]-[0164], wherein machine learning is performed based on a difference between a target sound quality index selected by a driver and an output sound quality index from the vehicle. See at least [0032], [0055], and [0094]-[0101], wherein the learning is done by an artificial intelligence.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to further modify Zhao with Jung’s technique of training an artificial intelligence to receive output sound and compare a sound quality index with a target sound quality index. It would have been obvious to modify because doing so enables optimized vehicle engine tone using artificial intelligence, as recognized by Jung (see at least [0006]-[0008]).
Regarding claim 8, Zhao, Kim, and Jung in combination teach all of the limitations of claim 7 as discussed above, and Zhao additionally teaches inputting, by a controller, vehicle information to an encoder; outputting sound through a decoder based on the third latent vector. See at least [0107], [n0052]-[n0054], and figure 3, wherein noise signals emitted by an engine of a vehicle are input into a parameter encoding layer (VAE) of an initial vehicle sound synthesis model. See at least [0062]-[0063], wherein the disclosed methods are performed by a processor. See at least [n0055]-[n0058] and figure 3, wherein the encoding from the first encoder is used by a decoding layer of the model to generate an output noise signal.
Zhao remains silent on a controller area network (CAN) signal and a vibration signal of the vehicle, the encoder being a third neural encoder, the decoder being a third neural decoder, generating a third latent vector for the CAN signal and vibration signal processed at the third neural encoder; and generating a sound quality index (SQI) from the output sound through a second neural network, wherein the third neural decoder is configured to use the second neural decoder, and wherein the second neural network is configured to use the first neural network.
Kim teaches the encoder being a third neural encoder, the decoder being a third neural decoder, generating a third latent vector for the signal processed at the third neural encoder. See at least [0110]-[0111] and figure 7, step S130, wherein multiple neural networks perform an encoding operations on the input data to generate latent data. See at least [0084], wherein the generated latent data comprises encoded code vectors. See at least [0003]-[0005] and [0052]-[0053], wherein the input data is audio signal data, and the encoders and decoders are performed by neural networks.
a second neural network, wherein the third neural decoder is configured to use the second neural decoder, and the second neural network is configured to use the first neural network. See at least [0055]-[0056] and [0059], second neural network NN2 configured to use data generated by the first neural network NN1. See at least [0161] and figure 4, wherein the neural networks NN1_1 to NN1_5, which include third neural decoder NN1_d3 and second neural decoder NN1_d2, are trained together.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to modify Zhao with Kim’s third neural encoder, third latent vector, third neural decoder, and second neural network. It would have been obvious to modify because doing so enables improvements in data transmission, in particular during compression and restoration of audio data, as recognized by Kim (see at least [0003]-[0007]).
Jung teaches a controller area network (CAN) signal and a vibration signal of the vehicle. See at least [0081], [0096], and [0147], wherein vehicle driving information 411 is input from a controller area network of the vehicle, along with engine characteristic information 310. See at least [0118], wherein engine characteristic information 310 includes vibration signal information.
and generating a sound quality index (SQI) from the output sound. See at least [0122], wherein a sound quality index is generated based on sound quality output results.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to further modify Zhao with Jung’s technique of collecting CAN signals and vibration signals of the vehicle and generating a sound quality index. It would have been obvious to modify because doing so enables optimized vehicle engine tone using artificial intelligence, as recognized by Jung (see at least [0006]-[0008]).
Regarding claim 9, Zhao, Kim, and Jung in combination teach all of the limitations of claim 8 as discussed above, and Zhao remains silent on further comprising training the third neural encoder by comparing the sound quality index output with a target sound quality index through the second neural network.
Kim teaches training the third neural encoder through the second neural network. See at least [0056] and [0059], wherein second neural network NN2 performs learning operations. See at least [0043], wherein learning operations are performed on different layers within the neural network.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to modify Zhao with Kim’s neural network. It would have been obvious to modify because doing so enables improvements in data transmission, in particular during compression and restoration of audio data, as recognized by Kim (see at least [0003]-[0007]).
Jung teaches further comprising training the artificial intelligence by comparing the sound quality index output with a target sound quality index. See at least [0070] and [0162]-[0164], wherein machine learning is performed based on a difference between a target sound quality index selected by a driver and an output sound quality index from the vehicle. See at least [0032], [0055], and [0094]-[0101], wherein the learning is done by an artificial intelligence.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to further modify Zhao with Jung’s technique of training an artificial intelligence to receive output sound and compare a sound quality index with a target sound quality index. It would have been obvious to modify because doing so enables optimized vehicle engine tone using artificial intelligence, as recognized by Jung (see at least [0006]-[0008]).
Regarding claim 10, Zhao, Kim, and Jung in combination teach all of the limitations of claim 7 as discussed above, and Zhao additionally teaches inputting, by a controller, vehicle information to an encoder; outputting sound through a decoder based on the third latent vector. See at least [0107], [n0052]-[n0054], and figure 3, wherein noise signals emitted by an engine of a vehicle are input into a parameter encoding layer (VAE) of an initial vehicle sound synthesis model. See at least [0062]-[0063], wherein the disclosed methods are performed by a processor. See at least [n0055]-[n0058] and figure 3, wherein the encoding from the first encoder is used by a decoding layer of the model to generate an output noise signal.
Zhao remains silent on a controller area network (CAN) signal and a vibration signal of the vehicle, the encoder being a third neural encoder, the decoder being a third neural decoder, generating a third latent vector for the CAN signal and vibration signal processed at the third neural encoder; and generating a sound quality index (SQI) from the output sound through a second neural network, wherein the third neural encoder is configured to use the second neural encoder, wherein the third neural decoder is configured to use the second neural decoder, and wherein the second neural network is configured to use the first neural network.
Kim teaches the encoder being a third neural encoder, the decoder being a third neural decoder, generating a third latent vector for the signal processed at the third neural encoder. See at least [0110]-[0111] and figure 7, step S130, wherein multiple neural networks perform an encoding operations on the input data to generate latent data. See at least [0084], wherein the generated latent data comprises encoded code vectors. See at least [0003]-[0005] and [0052]-[0053], wherein the input data is audio signal data, and the encoders and decoders are performed by neural networks.
a second neural network, wherein the third neural encoder is configured to use the second neural encoder, wherein the third neural decoder is configured to use the second neural decoder, and wherein the second neural network is configured to use the first neural network. See at least [0055]-[0056] and [0059], second neural network NN2 configured to use data generated by the first neural network NN1. See at least [0161] and figure 4, wherein the neural networks NN1_1 to NN1_5, which include third neural decoder NN1_d3, second neural encoder NN1_e2, and second neural decoder NN1_d2, are trained together.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to modify Zhao with Kim’s third neural encoder, third latent vector, third neural decoder, and second neural network. It would have been obvious to modify because doing so enables improvements in data transmission, in particular during compression and restoration of audio data, as recognized by Kim (see at least [0003]-[0007]).
Jung teaches a controller area network (CAN) signal and a vibration signal of the vehicle. See at least [0081], [0096], and [0147], wherein vehicle driving information 411 is input from a controller area network of the vehicle, along with engine characteristic information 310. See at least [0118], wherein engine characteristic information 310 includes vibration signal information.
and generating a sound quality index (SQI) from the output sound. See at least [0122], wherein a sound quality index is generated based on sound quality output results.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to further modify Zhao with Jung’s technique of collecting CAN signals and vibration signals of the vehicle and generating a sound quality index. It would have been obvious to modify because doing so enables optimized vehicle engine tone using artificial intelligence, as recognized by Jung (see at least [0006]-[0008]).
Regarding claim 11, Zhao, Kim, and Jung in combination teach all of the limitations of claim 10 as discussed above, and Zhao remains silent on further comprising: comparing the sound quality index output through the second neural network with a target sound quality index; and determining the CAN signal and vibration signal using an optimization algorithm.
Jung teaches further comprising: comparing the sound quality index output through the second neural network with a target sound quality index. See at least [0070] and [0162]-[0164], wherein machine learning is performed based on a difference between a target sound quality index selected by a driver and an output sound quality index from the vehicle. See at least [0032], [0055], and [0094]-[0101], wherein the learning is done by an artificial intelligence.
and determining the CAN signal and vibration signal using an optimization algorithm. See at least [0161]-[0164], wherein an order array is output using an optimization process. See at least [0102], wherein the order array refers to frequency and engine speed. See at least [0080]-[0081] and [0118] , wherein the engine speed and frequency are CAN and vibration signals.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to further modify Zhao with Jung’s technique of training an artificial intelligence to receive output sound and compare a sound quality index with a target sound quality index. It would have been obvious to modify because doing so enables optimized vehicle engine tone using artificial intelligence, as recognized by Jung (see at least [0006]-[0008]).
Regarding claim 12, Zhao teaches a method of controlling sound using a sound quality index-based generative artificial intelligence (AI), the method comprising: inputting, by a controller, sound generated. See at least [0107], [n0052]-[n0054], and figure 3, wherein noise signals emitted by an engine of a vehicle are input into a parameter encoding layer (VAE) of an initial vehicle sound synthesis model. See at least [0062]-[0063], wherein the disclosed methods are performed by a processor.
outputting sound through a first decoder based on the first latent vector. See at least [n0055]-[n0058] and figure 3, wherein the encoding from the first encoder is used by a decoding layer of the model to generate an output noise signal.
training the first neural encoder and the first neural decoder by comparing the output sound with a target sound. See at least [n0050], wherein the sound synthesis model, including the encoding layer and the decoding layer, is trained by comparing the output synthesis results with calibration sound results with target loss functions.
the sound being generated from an engine or a motor of the vehicle. See at least [0107], wherein the collected sound signals for input are generated from the engine through the exhaust system of the vehicle as the vehicle is driving.
inputting, by a controller, vehicle information to a second neural encoder. See at least [0093], [n0040]-[n0045], and figures 2-3, wherein vehicle parameters are input to a second encoding layer (paramVAE, separate from first encoding layer VAE).
and outputting sound through a second neural decoder based on the second latent vector, wherein the second neural decoder is configured to use the first neural decoder of which training is completed by the received sound, the output sound, and the target sound. See at least [n0045]-[n0048], wherein the output of the second encoder paramVAE is used by a decoding layer to output a sound signal. See at least [n0072], wherein the decoding layer of the trained vehicle sound synthesis model (second decoder, figure 2) is formed based on the decoding layer of the initial vehicle sound synthesis model (first decoder, figure 3). See at least [n0014] and [0040]-[0054], wherein the synthesis unit includes two decoding layers, one in the synthesis unit and one in a training subunit of the synthesis unit.
Zhao remains silent on the encoders being neural, generating a first latent vector for the sound in the first neural encoder, a controller area network (CAN) signal and vibration signal of the vehicle, and generating a second latent vector for the CAN signal and vibration signal processed in the second neural encoder.
Kim teaches the encoders being neural. See at least [0052]-[0053] and figure 4, wherein the encoders and decoders are performed by neural networks.
generating a first latent vector for the sound in the first neural encoder. See at least [0110]-[0111] and figure 7, step S130, wherein a first encoding neural network performs an encoding operation on the input data to generate first latent data. See at least [0084], wherein the generated latent data comprises encoded code vectors. See at least [0003]-[0005] and [0053], wherein the input data is audio signal data.
and generating a second latent vector for the CAN signal and vibration signal processed in the second neural encoder. See at least [0110]-[0111] and figure 7, step S130, wherein multiple neural networks perform an encoding operations on the input data to generate latent data. See at least [0084], wherein the generated latent data comprises encoded code vectors. See at least [0003]-[0005] and [0052]-[0053], wherein the input data is audio signal data, and the encoders and decoders are performed by neural networks.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to modify Zhao with Kim’s technique of neural encoders, and generating a first latent vector for the sound in the first neural encoder. It would have been obvious to modify because doing so enables improvements in data transmission, in particular during compression and restoration of audio data, as recognized by Kim (see at least [0003]-[0007]).
Jung teaches a controller area network (CAN) signal and a vibration signal of the vehicle. See at least [0081], [0096], and [0147], wherein vehicle driving information 411 is input from a controller area network of the vehicle, along with engine characteristic information 310. See at least [0118], wherein engine characteristic information 310 includes vibration signal information.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to further modify Zhao with Jung’s technique of collecting CAN signals and vibration signals of the vehicle. It would have been obvious to modify because doing so enables optimized vehicle engine tone using artificial intelligence, as recognized by Jung (see at least [0006]-[0008]).
Regarding claim 13, Zhao, Kim, and Jung in combination teach all of the limitations of claim 12 as discussed above, and Zhao additionally teaches wherein the controller is configured train the second neural encoder by comparing the sound output from the second neural decoder with the target sound. See at least [n0050] and [n0072], wherein the second encoder (paramVAE) is trained by comparing synthesized output sound with a target calibration sound signal.
Regarding claim 14, Zhao, Kim, and Jung in combination teach all of the limitations of claim 13 as discussed above, and Zhao remains silent on further comprising training a neural network configured to receive the output sound to thereby output a sound quality index (SQI) and compare the sound quality index with a target sound quality index.
Kim teaches a neural network. See at least [0055]-[0056], first neural network NN1.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to modify Zhao with Kim’s neural network. It would have been obvious to modify because doing so enables improvements in data transmission, in particular during compression and restoration of audio data, as recognized by Kim (see at least [0003]-[0007]).
Jung teaches further comprising training a first artificial intelligence configured to receive the output sound to thereby output a sound quality index (SQI) and compare the sound quality index with a target sound quality index. See at least [0070] and [0162]-[0164], wherein machine learning is performed based on a difference between a target sound quality index selected by a driver and an output sound quality index from the vehicle. See at least [0032], [0055], and [0094]-[0101], wherein the learning is done by an artificial intelligence.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to further modify Zhao with Jung’s technique of training an artificial intelligence to receive output sound and compare a sound quality index with a target sound quality index. It would have been obvious to modify because doing so enables optimized vehicle engine tone using artificial intelligence, as recognized by Jung (see at least [0006]-[0008]).
Regarding claim 15, Zhao teaches a method of controlling sound using a sound quality index-based generative artificial intelligence (AI), the method comprising: inputting, by a controller, sound generated. See at least [0107], [n0052]-[n0054], and figure 3, wherein noise signals emitted by an engine of a vehicle are input into a parameter encoding layer (VAE) of an initial vehicle sound synthesis model. See at least [0062]-[0063], wherein the disclosed methods are performed by a processor.
outputting sound through a first decoder based on the first latent vector. See at least [n0055]-[n0058] and figure 3, wherein the encoding from the first encoder is used by a decoding layer of the model to generate an output noise signal.
training the first neural encoder and the first neural decoder by comparing the output sound with a target sound. See at least [n0050], wherein the sound synthesis model, including the encoding layer and the decoding layer, is trained by comparing the output synthesis results with calibration sound results with target loss functions.
the sound being generated from an engine or a motor of the vehicle. See at least [0107], wherein the collected sound signals for input are generated from the engine through the exhaust system of the vehicle as the vehicle is driving.
inputting, by a controller, vehicle information to a second neural encoder. See at least [0093], [n0040]-[n0045], and figures 2-3, wherein vehicle parameters are input to a second encoding layer (paramVAE, separate from first encoding layer VAE).
outputting sound through a second neural decoder based on the second latent vector, wherein the second neural decoder is configured to use the first neural decoder of which training is completed by the received sound, the output sound, and the target sound. See at least [n0045]-[n0048], wherein the output of the second encoder paramVAE is used by a decoding layer to output a sound signal. See at least [n0072], wherein the decoding layer of the trained vehicle sound synthesis model (second decoder, figure 2) is formed based on the decoding layer of the initial vehicle sound synthesis model (first decoder, figure 3). See at least [n0014] and [0040]-[0054], wherein the synthesis unit includes two decoding layers, one in the synthesis unit and one in a training subunit of the synthesis unit.
inputting, by a controller, vehicle information to a neural encoder. See at least [0093], [n0040]-[n0045], and figures 2-3, wherein vehicle parameters are input to an encoder.
outputting sound through a neural decoder based on the latent vector. See at least [n0045]-[n0048], wherein the output of the encoder is used by a decoding layer to output a sound signal.
Zhao remains silent on the encoders being neural, generating a first latent vector for the sound in the first neural encoder, a controller area network (CAN) signal and vibration signal of the vehicle, generating a second latent vector for the CAN signal and vibration signal processed in the second neural encoder, a third neural encoder, generating, a third latent vector for the CAN signal vibration signal processed at the third neural encoder, a third neural decoder, wherein the third neural decoder is configured to use the second neural encoder; generating a sound quality index (SQI) from the output sound through a first neural network; training the first neural network by comparing the sound quality index output through the first neural network with a target sound quality index, a fourth neural encoder, wherein the fourth neural encoder is configured to use the third neural encoder, generating a fourth latent vector for the CAN signal vibration signal processed at the fourth neural encoder, a fourth neural decoder of which training is completed, wherein the fourth neural encoder is configured to use the third neural encoder, and generating a sound quality index (SQI) from the output sound through a second neural network, wherein the second neural network is configured to use the first neural network.
Kim teaches the encoders being neural. See at least [0052]-[0053] and figure 4, wherein the encoders and decoders are performed by neural networks.
generating a first latent vector for the sound in the first neural encoder. See at least [0110]-[0111] and figure 7, step S130, wherein a first encoding neural network performs an encoding operation on the input data to generate first latent data. See at least [0084], wherein the generated latent data comprises encoded code vectors. See at least [0003]-[0005] and [0053], wherein the input data is audio signal data.
generating a second latent vector for the CAN signal and vibration signal processed in the second neural encoder. See at least [0110]-[0111] and figure 7, step S130, wherein multiple neural networks perform an encoding operations on the input data to generate latent data. See at least [0084], wherein the generated latent data comprises encoded code vectors. See at least [0003]-[0005] and [0052]-[0053], wherein the input data is audio signal data, and the encoders and decoders are performed by neural networks.
the encoder being a third neural encoder, the decoder being a third neural decoder, generating a third latent vector for the signal processed at the third neural encoder. See at least [0110]-[0111] and figure 7, step S130, wherein multiple neural networks perform an encoding operations on the input data to generate latent data. See at least [0084], wherein the generated latent data comprises encoded code vectors. See at least [0003]-[0005] and [0052]-[0053], wherein the input data is audio signal data, and the encoders and decoders are performed by neural networks.
wherein the third neural encoder is configured to use the second neural encoder. See at least [0161] and figure 4, wherein the neural networks NN1_1 to NN1_5, which include third neural decoder NN1_d3, second neural encoder NN1_e2, and second neural decoder NN1_d2, are trained together.
a first neural network. See at least [0055]-[0056], first neural network NN1.
the encoder being a fourth neural encoder, the decoder being a fourth neural decoder, generating a fourth latent vector for the signal processed at the fourth neural encoder. See at least [0110]-[0111] and figure 7, step S130, wherein multiple neural networks perform an encoding operations on the input data to generate latent data. See at least [0084], wherein the generated latent data comprises encoded code vectors. See at least [0003]-[0005] and [0052]-[0053], wherein the input data is audio signal data, and the encoders and decoders are performed by neural networks.
wherein the fourth neural encoder is configured to use the third neural encoder, a fourth neural decoder of which training is completed,. See at least [0161] and figure 4, wherein the neural networks NN1_1 to NN1_5, which include fourth neural encoder NN1_e4, and third neural encoder NN1_e3, are trained together.
a second neural network, wherein the second neural network is configured to use the first neural network. See at least [0055]-[0056] and [0059], second neural network NN2 configured to use data generated by the first neural network NN1.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to modify Zhao with Kim’s technique of neural encoders, generating latent vectors at the respective neural encoders, and a first and second neural network. It would have been obvious to modify because doing so enables improvements in data transmission, in particular during compression and restoration of audio data, as recognized by Kim (see at least [0003]-[0007]).
Jung teaches a controller area network (CAN) signal and a vibration signal of the vehicle. See at least [0081], [0096], and [0147], wherein vehicle driving information 411 is input from a controller area network of the vehicle, along with engine characteristic information 310. See at least [0118], wherein engine characteristic information 310 includes vibration signal information.
generating a sound quality index (SQI) from the output sound through a first neural network; training the first neural network by comparing the sound quality index output through the first neural network with a target sound quality index. See at least [0070] and [0162]-[0164], wherein machine learning is performed based on a difference between a target sound quality index selected by a driver and an output sound quality index from the vehicle. See at least [0032], [0055], and [0094]-[0101], wherein the learning is done by an artificial intelligence.
and generating a sound quality index (SQI) from the output sound. See at least [0070] and [0162]-[0164], wherein machine learning is performed based on a difference between a target sound quality index selected by a driver and an output sound quality index from the vehicle. See at least [0032], [0055], and [0094]-[0101], wherein the learning is done by an artificial intelligence.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to further modify Zhao with Jung’s technique of collecting CAN signals and vibration signals of the vehicle. It would have been obvious to modify because doing so enables optimized vehicle engine tone using artificial intelligence, as recognized by Jung (see at least [0006]-[0008]).
Regarding claim 16, Zhao, Kim, and Jung in combination teach all of the limitations of claim 15 as discussed above, and Zhao remains silent on comprising tuning and additionally training the fourth neural encoder by comparing the sound quality index output through the second neural network with the target sound quality index.
Kim teaches comprising additionally training the fourth neural encoder through the second neural network. See at least [0056] and [0059], wherein second neural network NN2 performs learning operations. See at least [0043], wherein learning operations are performed on different layers within the neural network.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to modify Zhao with Kim’s neural network. It would have been obvious to modify because doing so enables improvements in data transmission, in particular during compression and restoration of audio data, as recognized by Kim (see at least [0003]-[0007]).
Jung teaches further comprising tuning the artificial intelligence by comparing the sound quality index output with a target sound quality index. See at least [0070] and [0162]-[0164], wherein machine learning is performed based on a difference between a target sound quality index selected by a driver and an output sound quality index from the vehicle. See at least [0032], [0055], and [0094]-[0101], wherein the learning is done by an artificial intelligence. See at least [0157]-[0162], wherein the learning comprises adjusting the frequency band and characteristics of the engine order and audio system.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to further modify Zhao with Jung’s technique of training an artificial intelligence to receive output sound and compare a sound quality index with a target sound quality index. It would have been obvious to modify because doing so enables optimized vehicle engine tone using artificial intelligence, as recognized by Jung (see at least [0006]-[0008]).
Regarding claim 17, Zhao, Kim, and Jung in combination teach all of the limitations of claim 16 as discussed above, and Zhao remains silent on comprising applying an optimization algorithm to the CAN signal and vibration signal of the vehicle by comparing the sound quality index output through the second neural network with the target sound quality index.
Jung teaches comprising applying an optimization algorithm to the CAN signal and vibration signal of the vehicle by comparing the sound quality index output through the second neural network with the target sound quality index. See at least [0161]-[0164], wherein an order array is output using an optimization process. See at least [0102], wherein the order array refers to frequency and engine speed. See at least [0080]-[0081] and [0118] , wherein the engine speed and frequency are CAN and vibration signals. See at least [0070] and [0162]-[0164], wherein the optimization process is performed based on a difference between a target sound quality index selected by a driver and an output sound quality index from the vehicle.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to further modify Zhao with Jung’s technique of training an artificial intelligence to receive output sound and compare a sound quality index with a target sound quality index. It would have been obvious to modify because doing so enables optimized vehicle engine tone using artificial intelligence, as recognized by Jung (see at least [0006]-[0008]).
Regarding claim 18, Zhao teaches a method of controlling sound using a sound quality index-based generative artificial intelligence (AI), the method comprising: inputting, by a controller, sound generated. See at least [0107], [n0052]-[n0054], and figure 3, wherein noise signals emitted by an engine of a vehicle are input into a parameter encoding layer (VAE) of an initial vehicle sound synthesis model. See at least [0062]-[0063], wherein the disclosed methods are performed by a processor.
outputting sound through a first decoder based on the first latent vector. See at least [n0055]-[n0058] and figure 3, wherein the encoding from the first encoder is used by a decoding layer of the model to generate an output noise signal.
training the first neural encoder and the first neural decoder by comparing the output sound with a target sound. See at least [n0050], wherein the sound synthesis model, including the encoding layer and the decoding layer, is trained by comparing the output synthesis results with calibration sound results with target loss functions.
the sound being generated from an engine or a motor of the vehicle. See at least [0107], wherein the collected sound signals for input are generated from the engine through the exhaust system of the vehicle as the vehicle is driving.
inputting, by a controller, vehicle information to a second neural encoder. See at least [0093], [n0040]-[n0045], and figures 2-3, wherein vehicle parameters are input to a second encoding layer (paramVAE, separate from first encoding layer VAE).
outputting sound through a second neural decoder based on the second latent vector, wherein the second neural decoder is configured to use the first neural decoder of which training is completed by the received sound, the output sound, and the target sound. See at least [n0045]-[n0048], wherein the output of the second encoder paramVAE is used by a decoding layer to output a sound signal. See at least [n0072], wherein the decoding layer of the trained vehicle sound synthesis model (second decoder, figure 2) is formed based on the decoding layer of the initial vehicle sound synthesis model (first decoder, figure 3). See at least [n0014] and [0040]-[0054], wherein the synthesis unit includes two decoding layers, one in the synthesis unit and one in a training subunit of the synthesis unit.
inputting, by a controller, vehicle information to a neural encoder. See at least [0093], [n0040]-[n0045], and figures 2-3, wherein vehicle parameters are input to an encoder.
outputting sound through a neural decoder based on the latent vector. See at least [n0045]-[n0048], wherein the output of the encoder is used by a decoding layer to output a sound signal.
Zhao remains silent on the encoders being neural, generating a first latent vector for the sound in the first neural encoder, a controller area network (CAN) signal and vibration signal of the vehicle, generating a second latent vector for the CAN signal and vibration signal processed in the second neural encoder, a third neural encoder, generating, a third latent vector for the CAN signal vibration signal processed at the third neural encoder, a third neural decoder, wherein the third neural decoder is configured to use the second neural encoder; generating a sound quality index (SQI) from the output sound through a first neural network; training the first neural network by comparing the sound quality index output through the first neural network with a target sound quality index, a fifth neural encoder, wherein the fifth neural encoder is configured to use the second neural encoder, generating a fifth latent vector for the CAN signal vibration signal processed at the fifth neural encoder, a fifth neural decoder, wherein the fifth neural decoder is configured to use the second neural decoder, and generating a sound quality index (SQI) from the output sound through a second neural network, wherein the second neural network is configured to use the first neural network.
Kim teaches the encoders being neural. See at least [0052]-[0053] and figure 4, wherein the encoders and decoders are performed by neural networks.
generating a first latent vector for the sound in the first neural encoder. See at least [0110]-[0111] and figure 7, step S130, wherein a first encoding neural network performs an encoding operation on the input data to generate first latent data. See at least [0084], wherein the generated latent data comprises encoded code vectors. See at least [0003]-[0005] and [0053], wherein the input data is audio signal data.
generating a second latent vector for the CAN signal and vibration signal processed in the second neural encoder. See at least [0110]-[0111] and figure 7, step S130, wherein multiple neural networks perform an encoding operations on the input data to generate latent data. See at least [0084], wherein the generated latent data comprises encoded code vectors. See at least [0003]-[0005] and [0052]-[0053], wherein the input data is audio signal data, and the encoders and decoders are performed by neural networks.
the encoder being a third neural encoder, the decoder being a third neural decoder, generating a third latent vector for the signal processed at the third neural encoder. See at least [0110]-[0111] and figure 7, step S130, wherein multiple neural networks perform an encoding operations on the input data to generate latent data. See at least [0084], wherein the generated latent data comprises encoded code vectors. See at least [0003]-[0005] and [0052]-[0053], wherein the input data is audio signal data, and the encoders and decoders are performed by neural networks.
wherein the third neural encoder is configured to use the second neural encoder. See at least [0161] and figure 4, wherein the neural networks NN1_1 to NN1_5, which include third neural decoder NN1_d3, second neural encoder NN1_e2, and second neural decoder NN1_d2, are trained together.
a first neural network. See at least [0055]-[0056], first neural network NN1.
the encoder being a fifth neural encoder, the decoder being a fifth neural decoder, generating a fifth latent vector for the signal processed at the fifth neural encoder. See at least [0110]-[0111] and figure 7, step S130, wherein multiple neural networks perform an encoding operations on the input data to generate latent data. See at least [0084], wherein the generated latent data comprises encoded code vectors. See at least [0003]-[0005] and [0052]-[0053], wherein the input data is audio signal data, and the encoders and decoders are performed by neural networks.
wherein the fifth neural encoder is configured to use the second neural encoder, wherein the fifth neural decoder is configured to use the second neural decoder. See at least [0161] and figure 4, wherein the neural networks NN1_1 to NN1_5, which include fifth neural encoder NN1_e5, fifth neural decoder NN1_d5, second neural encoder NN1_e2, and second neural decoder NN1_d2, are trained together.
a second neural network, wherein the second neural network is configured to use the first neural network. See at least [0055]-[0056] and [0059], second neural network NN2 configured to use data generated by the first neural network NN1.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to modify Zhao with Kim’s technique of neural encoders, generating latent vectors at the respective neural encoders, and a first and second neural network. It would have been obvious to modify because doing so enables improvements in data transmission, in particular during compression and restoration of audio data, as recognized by Kim (see at least [0003]-[0007]).
Jung teaches a controller area network (CAN) signal and a vibration signal of the vehicle. See at least [0081], [0096], and [0147], wherein vehicle driving information 411 is input from a controller area network of the vehicle, along with engine characteristic information 310. See at least [0118], wherein engine characteristic information 310 includes vibration signal information.
generating a sound quality index (SQI) from the output sound through a first neural network; training the first neural network by comparing the sound quality index output through the first neural network with a target sound quality index. See at least [0070] and [0162]-[0164], wherein machine learning is performed based on a difference between a target sound quality index selected by a driver and an output sound quality index from the vehicle. See at least [0032], [0055], and [0094]-[0101], wherein the learning is done by an artificial intelligence.
and generating a sound quality index (SQI) from the output sound. See at least [0070] and [0162]-[0164], wherein machine learning is performed based on a difference between a target sound quality index selected by a driver and an output sound quality index from the vehicle. See at least [0032], [0055], and [0094]-[0101], wherein the learning is done by an artificial intelligence.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to further modify Zhao with Jung’s technique of collecting CAN signals and vibration signals of the vehicle. It would have been obvious to modify because doing so enables optimized vehicle engine tone using artificial intelligence, as recognized by Jung (see at least [0006]-[0008]).
Regarding claim 19, Zhao, Kim, and Jung in combination teach all of the limitations of claim 18 as discussed above, and Zhao remains silent on comprising applying a Decision Engine to the CAN signal and vibration signal of the vehicle by comparing the sound quality index output through the third neural network with the target sound quality index.
Jung teaches comprising applying a Decision Engine to the CAN signal and vibration signal of the vehicle by comparing the sound quality index output through the third neural network with the target sound quality index. See at least [0161]-[0164], wherein an order array is output using an optimization process. See at least [0102], wherein the order array refers to frequency and engine speed. See at least [0080]-[0081] and [0118] , wherein the engine speed and frequency are CAN and vibration signals. See at least [0070] and [0162]-[0164], wherein the optimization process is performed based on a difference between a target sound quality index selected by a driver and an output sound quality index from the vehicle. The optimization is performed based on reinforcement by tone control operation unit 500.
One having ordinary skill in the art, before the effective filing date of the claimed invention, would have found it obvious to further modify Zhao with Jung’s technique of training an artificial intelligence to receive output sound and compare a sound quality index with a target sound quality index. It would have been obvious to modify because doing so enables optimized vehicle engine tone using artificial intelligence, as recognized by Jung (see at least [0006]-[0008]).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Selena M. Jin whose telephone number is (408)918-7588. The examiner can normally be reached Monday - Thursday and alternate Fridays, 7:30-4:30 PT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Faris Almatrahi can be reached at (313) 446-4821. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.M.J./ Examiner, Art Unit 3667
/FARIS S ALMATRAHI/ Supervisory Patent Examiner, Art Unit 3667