DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-5,12,14-15 and 17-18 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by US20220328025 (Ramirez), hereinafter US’025.
Regarding claim 1, US’025 discloses ‘A process of creating a controllable responsive model of an audio system (US’025, ¶[0019]; ¶[0020];¶¶[0022]-[0024], ¶¶[0064]-[0068]):, comprising:
generating a neural network that emulates a behavior of a reference audio system for at least two control settings of the reference audio system (US’025, ¶[0024]:” neural network may include a machine-learning model that can be tuned (e.g., trained) based on training input to approximate unknown functions…and generate outputs based on a plurality of inputs …”; ¶[0068]:neural network is based on training data), comprising:
performing for each control setting: receiving control position data designating a select control setting of the reference audio system as conditioning for the neural network (US’025, ¶[0023]:” the audio signal processing system 102 receives the audio input 100 from a user… the audio input 100 includes both unprocessed audio and a selection of an audio signal processing effect to be performed on the unprocessed audio”; ¶¶[0064]-[0068] providing control parameter information corresponding to user-selected control settings as conditioning information for the neural network during training;
communicating an input to the reference audio system and responsive thereto, capturing a target output of the reference audio system (US’025, ¶[0024]:”… a neural network can include a model of interconnected digital neurons that communicate and learn to approximate complex functions and generate outputs based on a plurality of inputs provided to the model”);
mapping parameters of the neural network such that, responsive to the input, a neural output resembles the target output of the reference audio system (US’025, ¶[0020];¶[0021];¶[0026]: the deep encoder estimates parameters, generates training output, updates; the model)
scoring by a loss function, a similarity of the neural network output compared to the target output of the reference audio system (US’025, ¶[0006]:”the deep encoder is trained using a loss function. This loss function is based on a comparison of the unprocessed audio input, the processed audio output”; ¶¶[0068]-[0070]:” the training system 710 trains the deep encoder 706 to learn to estimate parameters for one or more audio signal processing effects plugins using loss function 716. Loss function 716, as discussed above, compares processed audio generated by the audio effect module 708 to a target audio”);
and utilizing the similarity derived from the loss function to modify model parameters of the neural network to improve the scored similarity (US’025, ¶¶[0042]-[0044]; ¶¶[0068]-[0070]:” identify, generate, create, and/or determine training input and utilize the training input to train and fine-tune a neural network. … the training system … learn to estimate parameters … using loss function”);
and associating the neural network with a graphical user interface that is configured to enable a user to select a virtual control setting within the graphical user interface corresponding to a select one of the at least two control settings of the reference audio system such that the neural network models the reference audio system based upon the corresponding selected control setting (US’025, ¶[0023]:”a user may select an audio file including unprocessed audio in an application and be presented with an interface through which the user may input or select a type of audio signal processing effect, ¶[0062]:”graphical user interfaces ( or simply "user interfaces") that allow a user to view and interact with content”; ¶[0063]:user-selectable controls, menus, selectable options and interactive control elements).
Regarding claim 2, US’025 discloses ‘The process of claim 1, as discussed above.
wherein: the reference audio system includes at least two controls, each control capable of at least two control settings (US’025, Fig. 5, multiple settings, multiple adjustable parameter including Input Gain, Output Gain, Threshold, Ratio Frequency, Gain, Attack, Release, Knee, and Mix, each representing independently selectable control setting of reference audio system) ;
and generating the neural network that emulates the behavior of the reference audio system for at least two control settings further comprises: generating the neural network to emulate the behavior of the reference audio system in at least two control settings for each control of the reference audio system (US’025, ¶¶[0024]; ¶[0068]: training a deep neural network to estimate the operating parameter of the reference audio system using training data generated from the reference audio processing chain so that the network reproduces the behavior of the reference system over the parameter space).
Regarding claim 3, US’025 discloses ‘The process of claim 1, as discussed above.
wherein: generating the neural network that emulates the behavior of the reference audio system for at least two control settings, comprises: sampling a control space at a discrete collection of control positions for each control of the reference audio system (US’025, Figs. 1 and 5’ demonstrates that parameter values are generated and supplied to the plugin during training, ¶¶[0024] [0068], parameter values are estimated and use during training of the network);
and training the neural network to learn to generalize a continuum of control positions between a minimum discrete control position and a maximum discrete control position for each control (US’025, ¶¶[0068]-[0070], training the deep neural network using training examples so that the network predicts parameter values and reproduces the operation of the reference audio processor over its operating parameter range rather than memorizing individual training samples).
Regarding claim 4, US’025 discloses ‘The process of claim 1 as discussed above.
further comprising: training the neural network using variables grouped together as tuples, each tuple (US’025, Figs.1-5, illustrates relationship between input audio, parameter estimator, audio effects plugin, processed output and loss computation) including:
an input audio segment (US’025, Fig. 1, supplying input audio to the deep encoder during training);
a recording of the target output of the reference audio system responsive to the input audio segment, the recording used as the target for the neural output (Fig. 1 and Fig. 5, generating processed audio from the reference audio effects plugin which serves as the target output during supervised learning; ¶[0022]:” raw audio through the deep encoder to generate parameters for one or more audio signal processing effects plugins, transforms the unprocessed, raw audio into a produced audio recording”);
and control values describing the control settings of the reference audio system (US’025, Fig. 5 ¶¶[0062]-[0068]:” the training system 710 trains the deep encoder 706 to learn to estimate parameters for one or more audio signal processing effects plugins”, parameter values estimated for the audio effects plugin corresponding to the operating controls of the reference audio processor).
Regarding claim 5, US’025 discloses ‘The process of claim 1, as discussed above.
wherein, the reference audio system includes a control capable of at least two control settings, wherein: the control comprises a select one of: a potentiometer or encoder, where the least two control settings span a range of successive control values (US’025, Fig.5, continuously variable plugin parameters including gain, threshold, ratio, frequency, attack, release, knee, mix, and similar user-adjustable parameters represented graphically within the interface);
a switch having at least two switch positions (US’025, ¶[0023], ¶¶[0062]-[0063], selectable processing options and plugin selections through the graphical interface);
or a time varying control that changes over time at a slower rate than a range of intended frequencies at the input of the neural network (US’025, ¶[0066]:” deep encoder processes frames of the unprocessed audio input individually or in a batch, and parameters are estimated … parameters may include any parameters … threshold, makeup gain, ratio, frequency splits, input gain, output gain”,¶[0068], estimating and applying parameter values during operation of neural-network based audio processor. Parameter values may change during processing).
Regarding claim 12, US’025 discloses ‘The process of claim 1, as discussed above.
wherein: mapping parameters of the neural network such that, responsive to the input, a neural output resembles the target output of the reference audio system (US’025, ¶¶[0065]-[0068], Figs. 2-5, training a neural network model whose parameters are optimized so the model output reproduces the output of the target), comprises:
accepting an input audio segment (x) (US’025, Fig. 3, ¶[0039] ¶[0041], providing an audio input signal to the neural network for processing) and a control value (c) describing the control settings of the reference audio system (US’025, Fig. 3, ¶[0023], ¶[0040], ¶[0062]), conditioning the neural network on user-selected control values) that is parameterized by network parameters θ (US’025, ¶[0068]: learn to estimate parameters … to minimize the loss” ¶[0027]:” Weights associated with each output node may determine the parameter values used by one or more audio signal processing”, learned weights are network parameters);
and scoring by the loss function, the similarity of the neural network output compared to the target output of the reference audio system (US’025, ¶¶0038]-[0044]:”the loss function receives the processed audio and a target audio...”¶[0039]: “the loss function…is used to calculate the gradients of the processed audio..”), comprises:
measuring a discrepancy between the neural network output compared to the target output of the reference audio system (US’025, ¶[0042], ¶[0043], ¶[0066]) using standard supervised learning (US’025, ¶[0065]) with a stochastic gradient descent (US’025, ¶[0052]) to learn the network parameters θ that minimize the average loss over a corresponding training set (US’025, ¶[0068]: learn to estimate parameters … to minimize the loss”).
Regarding claim 14, US’025 discloses ‘A system defining a controllable responsive model of an audio system (US’025, ¶[0019]:”an audio signal processing system that uses machine learning to perform audio signal processing”), comprising: a first processing system configuration operatively programmed (US’025, ¶[0023]:” audio signal processing system 102 receives the audio input 100 from a user via a computing device”) to generate a neural network that emulates a behavior of a reference audio system for at least two control settings of the reference audio system, the first processing system configuration programmed to perform, for each control setting, operations that:
receive control position data designating a select control setting of the reference audio system as conditioning for the neural network;
communicate an input to the reference audio system and responsive thereto, capture a target output of the reference audio system;
map parameters of the neural network such that, responsive to the input, a neural output resembles the target output of the reference audio system;
score by a loss function, a similarity of the neural network output compared to the target output of the reference audio system;
and utilize the similarity derived from the loss function to modify model parameters of the neural network to improve the scored similarity;
and a graphical user interface associated with a second processing system configuration, the graphical user interface associated with the neural network to enable a user to select a virtual control setting within the graphical user interface corresponding to a select one of the at least two control settings of the reference audio system such that the neural network models the reference audio system based upon the corresponding selected control setting. (Claim 14 corresponds to claim 1)
Regarding claim 15, US’025 discloses ‘The system of claim 14, as discussed above.
US’025 further discloses wherein: the first processing system configuration is implemented by a computer system (US’025, ¶¶[0065]-[0068]);
and the second processing system configuration (US’025, ¶¶[0062]-[0063]:, application executing the trained model with GUI) is implemented in a second computer system that is different from the first computer system (US’025, ¶¶[0065]-[0068]:” audio effects module 708 can receive or retrieve unprocessed audio and parameters associated with the one or more audio signal processing effects plugins 714A-714N from a computing device”).
Regarding claim 17, US’025 discloses ‘The system of claim 14, as discussed above.
wherein: the reference audio system includes at least two controls, each control capable of at least two control settings;
and the first processing system configuration generates the neural network that emulates the behavior of the reference audio system for at least two control settings by generating the neural network to emulate the behavior of the reference audio system in at least two control settings for each control of the reference audio system. (Claim 17 corresponds to claim 2)
Regarding claim 18, US’025 discloses ‘The process of claim 14, as discussed above.
wherein: the reference audio system generates the neural network that emulates the behavior of the reference audio system by executing code that:
samples of a control space at a discrete collection of control positions for each control of the reference audio system;
and trains the neural network to learn to generalize a continuum of control positions between a minimum discrete control position and a maximum discrete control position for each control. (Claim 18 corresponds to claim 3)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 6-11 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over US’025, in view of US10073887 (Chehreghani), hereinafter US’887.
Regarding claim 6, US’025 discloses ‘The process of claim 1 as discussed above.
US’025 further discloses training a neural network ‘to determine the at least two control settings of the reference audio system’ (US’025, ¶[0024};¶[0028])
US’025 further discloses optimizing the training process using a delay-variant loss function that searches within a small time window to determine temporal alignment between the neural-network output and the processed reference output (US’025, ¶[0043]:”The delay-invariant loss function can find the best matching time point between the processed audio output across a small-time window and then optimize the loss”).
US’025 does not expressly disclose ‘further comprising: combining a random sampling procedure with an optimized sorting approach; generating all measured control configurations ahead of time into a list; and utilizing a sorting of the list is designed to minimize the overall distance travelled.
US’887 discloses further comprising: combining a random sampling procedure with an optimized sorting approach to determine the at least two control settings of the reference audio system (US’887, col. 6, lines 6-9:” for each unselected node at time T, the minimum pairwise distance to any one of a current subset 80 of objects (nodes) (FIGS. 3 and 4). These 45 minimum pairwise distances may be stored in a distance vector”; col. 14, lines 56-59:” Kruskal algorithm (Kruskal, "On the Shortest Spanning Subtree of a Graph and the Traveling Salesman Problem," Proc. Am. Mathematical Soc., 7, pp. 48-50, 1956) and the Prim algorithm constructing a graph of nodes using pairwise distances, optimizing the ordering of sampled configurations using graph-based optimization);
generating all measured control configurations ahead of time into a list (US’887, col. 7, lines 41-45:” system 40 uses an iterative approach for selecting nodes for the Minimax K-NN set 44 of dataset objects… component 72 identifies, for each unselected node at time T, the minimum pairwise distance to any one of a current subset 80 of objects (nodes) (FIGS. 3 and 4)”, candidate nodes collectively before constructing the graph from the available nodes);
and utilizing a sorting of the list is designed to minimize the overall distance travelled (US’887, col. 14, lines 56-61:” Kruskal … Shortest Sparming Subtree of a Graph and the Traveling Salesman Problem … and the Prim algorithm … at each step, picks the pair of candidate trees that have the minimal distance”, selecting minimum-distance connections and using minimum spanning tree; sorting or ordering nodes to minimize travel through the parameter space).
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to modify the training process of US’025 to order sampled control configurations using the graph optimization of US’887 because US’887 teaches ordering nodes based on pairwise distances to minimize overall traversal distance. Applying this optimization technique would improve the efficiency of traversing the sample control configurations used to train the neural network.
Regarding claim 7, US’025 (in view of US’887) discloses ‘The process of claim 6 as discussed above.
‘further comprising: designating that the process starts and finishes recording with all controls at zero;
US’025 discloses establishing predetermined control settings of the reference audio system for collecting training date before optimization of the neural network (US’025, ¶[0023],¶[0024],[0068]:”the training system 710 trains the deep encoder 706 to learn to estimate parameters for one or more audio signal processing effects plugins”)
US’887 discloses initializing the optimization by selecting an initial node and initializing the distance vector before the iterative traversal begins (US’887, col. 11, line 58:Initialize vector dist…” and col. 7, lines 58-59: “the iterative process sequentially adds a node…”, predetermined initial reference control configuration from which optimization begins to which it may return upon completion)
and finding the optimal path through random samples as a traveling salesman problem (TSP) (US’887, col. 14, lines 56-61:” Kruskal … Shortest Sparming Subtree of a Graph and the Traveling Salesman Problem … and the Prim algorithm … at each step, picks the pair of candidate trees that have the minimal distance”).
Regarding claim 8, US’025 (in view of US’887) discloses ‘The process of claim 6 as discussed above.
US’025 (in view of US’887) further discloses ‘further comprising: generating nodes as a list of knob-positions-to-visit (US’025, ¶[0024], control configurations) by random sampling, each knob-position-to-visit (US’025, ¶[0024]) corresponding to a select node (US’887, col 7 , lines 46-59:” The subset 80 of objects consists of node v and all nodes which have so far been selected to be in the K-NN set … The extension component 74 extends the partial set … by adding…one of the unselected nodes … The iterative process thus sequentially adds a node to set 84”; configurations are represented as nodes in the graph);
creating a distance matrix that is formed pair-wise for the nodes (US’887, col. 12, lines 14-15:”The measurements D may be initially 15 stored as a matrix”; col. 3, lines 59-60“computing a pairwise distance between a test object and the dataset object”);
and approximating a traveling salesman solution on the distance matrix (US’887, col. 14, lines 56-61:” Kruskal … Shortest Sparming Subtree of a Graph and the Traveling Salesman Problem … and the Prim algorithm … at each step, picks the pair of candidate trees that have the minimal distance”).
Regarding claim 9, US’025 (in view of US’887) discloses ‘The process of claim 8 as discussed above.
US’025 (in view of US’887) further discloses ‘further comprising: selecting a starting node (US’887, col. 11, lines 59-60: Initialize vector dist by the distance of v to each object… NP(v) = [ ]”, initializing process with a starting node);
and visiting a nearest unvisited node until all nodes are visited (US’887, col. 12, lines 48-54, “The node with the shortest pairwise distance …is added to the set … At the next iteration, the unselected node with the shortest pairwise distance … The next node to be added… and so forth until K nodes have been added”, repeatedly selecting the nearest unvisited node).
Regarding claim 10, US’025 (in view of US’887) discloses ‘The process of claim 9, as discussed above.
‘wherein: selecting a starting node comprises all knobs at zero.
US’887 discloses initializing the optimization by selecting an initial node before the beginning the optimization process (US’887, col. 11, line 58:Initialize vector dist…” selecting the initial node from which the graph traversal proceeds, Applicant’s Specification explains that “all knobs at zero” may alternatively be “a minimal position that still allows the reference audio system to produce sound” (Spec. ¶[0132])
Regarding claim 11, US’025 (in view of US’887) discloses ‘The process of claim 9, as discussed above.
US’025 (in view of US’887) further discloses ‘wherein: visiting the nearest unvisited node comprises looking up the nearest unvisited node in the distance matrix (US’887, col. 8, lines 65-67:”At S110, a distance vector 82 is computed …The distance vector includes, for each unselected node…”; col. 11, lines 48-50:”updates dist by checking if an unselected node has a smaller pairwise distance to the new .NP r(v).Thus dist is updated”); col. 12, lines 14-15:” The measurements D may be initially stored as a matrix”).
Regarding claim 19, US’025 discloses ‘The system of claim 14, as discussed above.
wherein the first processing system configuration further: combines a random sampling procedure with an optimized sorting approach to determine the at least two control settings of the reference audio system;
generates all measured control configurations ahead of time into a list;
and utilizes a sorting of the list that is designed to minimize the overall distance travelled. (Claim 19 corresponds to claim 6)
Regarding claim 20, US’025 discloses ‘The process of claim 14, as discussed above.
wherein the first processing system configuration further: generates nodes as a list of knob-positions-to-visit by random sampling, each knob-position-to-visit corresponding to a select node;
creates a distance matrix that is formed pair-wise for the nodes;
and approximates a traveling salesman solution on the distance matrix. (Claim 20 corresponds to claim 8)
Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over US’025, in view of US9589549 (McKay), hereinafter US’549.
Regarding claim 13, US’025 discloses ‘The process of claim 1, as discussed above.
‘further comprising adjusting the reference audio system to each control setting using a robot that physically connects to controls of the reference audio system.
US’025 further discloses generating training data by adjusting the reference audio system to each control setting of the reference audio system.
US’025 does not expressly disclose ‘using a robot that physically connects to controls.
However US’549 discloses ‘further comprising adjusting the reference audio system to each control setting (US’549, col. 3, lines 11-15:”a motor … coupling 2 coupled to the motor and the potentiometer for changing the position of the manual potentiometer”) using a robot (US’549, col. 2, lines 1-4:” robotic mechanism will utilize a widely used pivot-action foot pedal …to control this knob-adjusting device”) that physically connects to controls of the reference audio system (US549, col. 4, lines 27-28:” this product can be attached to the potentiometer of any pedal”)
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of the claim invention to modify the training process of US’025 by using US’549’s robotic control mechanism to automatically adjust the control settings of the reference audio system because McKay teaches a robotic mechanism having a motor and hardware coupling that physically attaches to third-party potentiometers to automatically change control settings, thereby improving efficiency and consistency when collecting training data for multiple control settings.
Claim 16 is rejected under 35 U.S.C. 103 as being unpatentable over US’025, in view of US20100175543 (Robertson), hereinafter US’543.
Regarding claim 16, US’025 discloses ‘The system of claim 15, as discussed above.
US’025 discloses ‘wherein: The second computer system US’025, ¶¶[0062]-[0068]) comprises’ implementing the neural-network model ‘for performing using the neural network at the user selected control setting (US’025, “¶[0024]:” neural network may include a machine-learning model that can be tuned (e.g., trained) based on training input to approximate unknown functions”).
US’025 does not expressly disclose ‘a dedicated hardware guitar effects processor that enables an instrument to be plugged directly therein.
However, US’543 discloses ‘a dedicated hardware guitar effects processor (US’543, ¶¶[0007]-[0008]:” software application … amplifies and processes electrical guitar signals”; ¶[0028]:” The signal received … is processed in real time by the digital signal processing and guitar effects software block 150”) that enables an instrument to be plugged directly therein (US’543, ¶¶[0038]-[0042]:” The audio coupling cable 520 further comprises a male mono plug 550 to provide a connection to the guitar”).
It would have been obvious to one of ordinary skill in the art prior to the effective filing date of claimed invention to implement the trained neural-network model of US’025 using the guitar processing hardware interface of US’543 because US’543 teaches a dedicated guitar-effects processing system configured to receive a guitar input directly and perform real-time digital signal processing. Incorporating US’543’s hardware interface into US’025 would have enabled direct connection of an instrument to execute the trained neural-network model while maintaining real-time performance and efficient usability for musicians.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. US20220180766 teaches adaptive and interactive teaching of playing a musical instrument.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NICOLE K GILLESPIE whose telephone number is (571)482-4187. The examiner can normally be reached Monday-Friday 7:30-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Dedei K Hammond can be reached at (571)270-3819. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NICOLE K GILLESPIE/Examiner, Art Unit 2837
/DEDEI K HAMMOND/Supervisory Patent Examiner, Art Unit 2837