Notice of Pre-AIA or AIA Status
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
2. This action is in response to the original filing on 09/16/2024. Claims 1-12 are pending and have been considered below.
Information Disclosure Statement
3. The information disclosure statement (IDS(s)) submitted on 09/16/2024, 12/04/2024, 07/09/2025 is/are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
4. 35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-12 are rejected under 35 U.S.C. 101 because the claimed invention is directed to the abstract idea without significantly more.
Step 1, the claims are directed to a process, machine, and manufacture.
Step 2A Prong 1, Claims 1, 11, 12 recites, in part
generating, fingering information indicative of fingering by processing the acquired input information (Mental processes, a music instructor could observe a player’s finger placement on a fretboard, listen to the resulting sound, use previously learned knowledge concerning finger positions and sounds, and make a judgement as to which fingering is being used).
Step 2A Prong 2, this judicial exception is not integrated into a practical application.
The additional elements:
a computer-implemented method for processing information that is executable by a computer system (mere instructions to apply the exception using a generic computer component).
acquiring, input information including: finger information relating to fingers of a user playing a string instrument and an image of a fretboard of the string instrument; and sound information of sound of the string instrument played by the user (mere data gathering and recited at a high level of generality, and thus are insignificant extra-solution activity).
using at least one generation model that learns a relationship between training input information and training fingering information (merely uses a computer as a tool to perform an abstract idea).
Step 2B, the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception, either alone or in combination.
The additional elements:
a computer-implemented method for processing information that is executable by a computer system (mere instructions to apply the exception using a generic computer component).
acquiring, input information including: finger information relating to fingers of a user playing a string instrument and an image of a fretboard of the string instrument; and sound information of sound of the string instrument played by the user (mere data gathering and recited at a high level of generality, and thus are insignificant extra-solution activity).
using at least one generation model that learns a relationship between training input information and training fingering information (merely uses a computer as a tool to perform an abstract idea).
Claims 2-10 provide further limitations to the abstract idea (Mathematical concepts and/or Mental processes) as rejected in claim 1, however, they do not disclose any additional elements that would amount to a practical application or significantly more than an abstract idea (data gathering/insignificant extra-solution activity and/or generic computer component).
Claim Rejections – 35 USC § 103
5. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
6. Claims 1, 7, 11, and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Avitabile et al. (U.S. Patent Application Pub. No. US 20150027297 A1) in view of Lee et al. (Observing Pianist Accuracy and Form with Computer Vision, IEEE, published 2019, pages 1505-1513).
Claim 1: Avitabile teaches a computer-implemented method for processing information that is executable by a computer system (i.e. FIG. 4 shows an embodiment of the tablet 100 in more detail. It can be seen that the tablet 100 comprises a central processing unit (CPU) 400, a memory 402 and a storage unit 414; para. [0037]), the method comprising:
acquiring, by the computer system, input information including (i.e. The tablet 100 also comprises a microphone 404 for receiving audio data, a camera 406 for receiving image data and a receiver 408 for receiving the force data transmitted by the transmitter 302 of the finger tip sensor 102. As will be discussed later, the musical tuition software application may use at least one of the audio data, image data and force data received by the tablet for providing feedback to the user on their performance with the musical instrument; para. [0038, 0040, 0081]):
finger information relating to fingers of a user playing a string instrument (i.e. Column 606 shows the position of the finger of the user on which the finger tip sensor 102 is worn. The position has two components, these being the one or more strings of the guitar which are pressed down by the user 610 and the number of the fret onto which the one or more strings are pressed 612 … The position of the finger on which the finger tip sensor 102 is worn is detected from the image data ID captured by the camera 406 of the tablet 100 using any suitable image analysis method; para. [0045, 0047, 0073]) and an image of a fretboard of the string instrument (i.e. On the calibration screen 617, there is shown a live image 615 of the user's guitar which is being captured by the camera 406 of the tablet 100. The user has positioned the tablet and guitar so that the fret board 620 of the guitar, including the first fret 621, can be seen by the camera; para. [0060, 0062, 0120]); and
sound information of sound of the string instrument played by the user (i.e. The tablet 100 also comprises a microphone 404 for receiving audio data; para. [0038]); and
generating, by the computer system, fingering information indicative of fingering (i.e. Through the detection and numbering of the strings 625 and fret bars 626 and the detection of the coloured finger tip sensors 102A-D, the correlation function engine 500 is able to determine the position of each of the finger tip sensors with respect to a string and fret bar; para. [0072-0077]) by processing the acquired input information (i.e. It is noted that each of audio data AD, force data FD and/or image data ID captured by the tablet 100 can be used in combination in order to ensure that the table 600 is as accurate as possible … processing is performed on the live image 930 and on force data FD generated by the finger tip sensors 102A-D and/or audio data AD detected by the microphone 404 so as to detect the positions of the user's fingers with respect to the strings 624 and frets 626 of the guitar; para. [0052, 0056, 0080, 0120]) using at least one generation model (i.e. The software application includes a correlation function engine 500 which converts the audio data AD, force data FD and/or image data ID received by the tablet 100 during a musical instrument performance into a correlation function CF; para. [0040]).
Avitabile does not explicitly teach model that learns a relationship between training input information and training fingering information.
However, Lee teaches generating, by the computer system, fingering information indicative of fingering (i.e. We present a first step towards developing an interactive piano tutoring system that can observe a student playing the piano and give feedback about hand movements and musical accuracy. In particular, we have two primary aims: 1) to determine which notes on a piano are being played at any moment in time, 2) to identify which finger is pressing each note. We introduce a novel two-stream convolutional neural network that takes video and audio inputs together for detecting pressed notes and finger presses. We formulate our two problems in terms of multi-task learning and extend a state-of-the-art object detection model to incorporate both audio and visual features. In addition, we introduce a novel finger identification solution based on pressed piano note information. We experimentally confirm that our approach is able to detect pressed piano keys and the piano player's fingers with a high accuracy; abs) by processing the acquired input information using at least one generation model (i.e. a novel two-stream convolutional neural network that takes video and audio inputs together for detecting pressed notes and finger presses; abs) that learns a relationship between training input information (i.e. We created new datasets of people playing the piano for training and testing our techniques. Figure 2 shows the pipeline that we used for generating our dataset given three different input files recorded while people played the piano: videos, MIDI files, and music scores. First, we extracted image frames and audio from the input video, and then applied the pre-trained hand detector on image frames to obtain bounding boxes of fingers. For the audio stream, we extracted MFCC features with 100 ms windows, and converted these into multi-channel images based on the first and second order derivatives, as described above. We extracted keypress information from MIDI synchronized with the video to create keypress ground truth labels; Section 4.1) and training fingering information (i.e. We then manually labeled each bounding box with the finger num-ber(s) that are currently pressing keys, and train the proposed network on this dataset … We collected 65 minutes of video also at 60 fps, and annotated them frame by frame for both notes and fingering using a combination of MIDI and manual labeling; Section 3.2.2, 4.1).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Avitabile to include the feature of Lee. One would have been motivated to make this modification because it improves the accuracy and robustness of determination of the user’s fingering, particularly where the image or audio information alone is ambiguous.
Claim 7: Avitabile and Lee teach the method for processing information according to claim 1. Avitabile further teaches comprising: generating, by the computer system, content based on the sound information and the fingering information (i.e. FIG. 6A shows an example of a correlation function CF according to an embodiment of the present disclosure. The correlation function CF comprises a table 600 recording various parameters of the musical performance which have been obtained from the audio data AD, force data FD and/or image data ID received by the tablet 100. The table 600 includes four columns 602, 604, 606 and, 608 and is a table which relates to a single finger tip sensor 102 worn by the user. A further column 620 in the table 600 is included that relates to the sound received from the microphone 404 of the tablet. A separate table 600 will be generated for each finger tip sensor 102 worn by the user, and the combination of all the tables gives the correlation function CF; para. [0042]).
Claim 11 is similar in scope to Claim 1 and is rejected under a similar rationale.
Avitabile teaches an information processing system comprising: at least one memory storing a program; at least one processor configured to execute the program to (i.e. FIG. 4 shows an embodiment of the tablet 100 in more detail. It can be seen that the tablet 100 comprises a central processing unit (CPU) 400, a memory 402 and a storage unit 414. The storage unit may comprise any suitable medium for storing electronic data, such as a flash memory unit or hard disk drive (HDD). The software application for providing the interactive musical tuition to the user is stored in the storage unit 414 and is executed by the CPU 400 when an instruction is received from the user to run the software application; para. [0037]).
Claim 12 is similar in scope to Claim 1 and is rejected under a similar rationale.
Avitabile teaches a computer-readable non-transitory storage medium for storing a program executable by a computer system to execute a method of (i.e. FIG. 4 shows an embodiment of the tablet 100 in more detail. It can be seen that the tablet 100 comprises a central processing unit (CPU) 400, a memory 402 and a storage unit 414. The storage unit may comprise any suitable medium for storing electronic data, such as a flash memory unit or hard disk drive (HDD). The software application for providing the interactive musical tuition to the user is stored in the storage unit 414 and is executed by the CPU 400 when an instruction is received from the user to run the software application; para. [0037]).
7. Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Avitabile in view of Lee, and further in view of Perez-Carrillo et al. (Estimation of Guitar Fingering and Plucking Controls Based on Multimodal Analysis of Motion, Audio and Musical Score, published 2016, pages 71-87).
Claim 2: Avitabile and Lee teach the method for processing information according to claim 1. Avitabile does not explicitly teach detecting one or more onsets of the string instrument, wherein acquisition of the input information and generation of the fingering information are executed for each of the one or more onsets.
However, Perez-Carrillo teaches detecting one or more onsets of the string instrument (i.e. The plucking instants are determined from the audio by means of a note onset detection algorithm [7] followed by an alignment and match to the note start times in the musical score; Section 4.1), wherein acquisition of the input information (i.e. The procedure for the parameter estimation proposed in this work is shown in Fig. 1. The algorithm starts with (1) the estimation of the plucking instants and the pitch by onset analysis of the audio signal followed by (2) an alignment to the score. At each plucking instant (3) the distances from the fingers to the strings are computed and (4) the possible combination of fret and string is estimated from the pitch … Distance from the markers in the fingers to the strings is computed as the average distance in a small window around the note onsets; Section 1, 4.2) and generation of the fingering information are executed for each of the one or more onsets (i.e. The procedure for the parameter estimation proposed in this work is shown in Fig. 1. The algorithm starts with (1) the estimation of the plucking instants and the pitch by onset analysis of the audio signal followed by (2) an alignment to the score. At each plucking instant (3) the distances from the fingers to the strings are computed and (4) the possible combination of fret and string is estimated from the pitch … Distance from the markers in the fingers to the strings is computed as the average distance in a small window around the note onsets; Section 1, 4.2).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Avitabile and Lee to include the feature of Perez-Carrillo. One would have been motivated to make this modification because it permits reliable detection of note onsets and uses the detected plucking instants as temporal reference points for multimodal fingering analysis.
8. Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Avitabile in view of Lee, and further in view of Iwase et al. (U.S. Patent Application Pub. No. US 20110023691 A1)
Claim 3: Avitabile and Lee teach the method for processing information according to claim 1. Avitabile does not explicitly teach generating, based on the fingering information, musical score information indicative of a musical score for playing the string instrument by the user.
However, Iwase teaches generating, by the computer system, based on the fingering information (i.e. The musical performance information acquiring section 23 acquires fingering information indicating the positions of the fingers of the performer; para. [0061]), musical score information indicative of a musical score (i.e. The output audio signal and musical performance information are used for score display or the like. For example, a score is displayed on the monitor on the basis of the note number included in the musical performance information, and musical sound is emitted simultaneously, such that the score can be used as a teaching material for training; para. [0248]) for playing the string instrument by the user (i.e. a score may be generated on the basis of the musical performance information. Therefore, a composer can generate a score by playing only the guitar 1; para. [0089]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Avitabile and Lee to include the feature of Iwase. One would have been motivated to make this modification because a composer can generate a score simply by playing the guitar, thereby avoiding complicated manual transcription.
9. Claims 4 and 6 are rejected under 35 U.S.C. 103 as being unpatentable over Avitabile in view of Lee, and further in view of Villa et al. (U.S. Patent Application Pub. No. US 20110045907 A1)
Claim 4: Avitabile and Lee teach the method for processing information according to claim 1. Avitabile does not explicitly teach showing on a display, by the computer system, a reference image representative of: a virtual player with fingering indicated by the fingering information; and a virtual string instrument played with the fingering.
However, Villa teaches showing on a display, by the computer system (i.e. the console 106 provides video output to a display 112; para. [0033]), a reference image representative of: a virtual player (i.e. the image comprises an avatar playing a representation of the musical instrument. In some embodiments, the image comprises an actual picture or likeness of the user playing the musical instrument. The image may then be displayed on a display; para. [0053]) with fingering indicated by the fingering information (i.e. the console deciphers what user manipulation of the musical instrument, such as the fingering, blowing, or other technique, is required to create the sound produced by the musical instrument. The console then applies the deciphered user manipulation of the musical instrument to an image, such as an avatar. In this scenario the avatar mimics the deciphered user manipulation of the musical instrument on a representation of the musical instrument; para. [0067]); and a virtual string instrument played with the fingering (i.e. The console 106 receives the signal representative of the sound produced by the guitar and deciphers what note or chord is being played. The console 106 then deciphers what musical instrument fingering is required to create the sound produced by the guitar. The avatar is then caused to mimic the deciphered musical instrument fingering on the representation of the guitar. In this way the image 702 is caused to be responsive to the signal representative of sound produced by the guitar; para. [0068]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Avitabile and Lee to include the feature of Villa. One would have been motivated to make this modification because the detected fingering could be presented visually to the user in an intuitive form showing how the fingering is performed on the instrument.
Claim 6: Avitabile, Lee, and Villa teach the method for processing information according to claim 4. Avitabile does not explicitly teach wherein the displaying the reference image includes: transmitting, by the computer system, image data indicative of the reference image to the terminal apparatus via a network: and displaying, by the terminal apparatus, the reference image transmitted from the computer system, on the display of the terminal apparatus.
However, Villa further teaches wherein the display is included in a terminal apparatus (i.e. The system 900 may be coupled to, or integrated with, any of the other components described herein, such as the display 112, amplifier 108, speaker 110, receiver 120, and/or microphone 204; para. [0078]), and wherein the displaying the reference image includes (i.e. the image comprises an avatar playing a representation of the musical instrument. In some embodiments, the image comprises an actual picture or likeness of the user playing the musical instrument. The image may then be displayed on a display; para. [0053]): transmitting, by the computer system, image data indicative of the reference image to the terminal apparatus via a network (i.e. The console 106 generates an image 502 that is representative of the user 102 playing the guitar 104. The image 502 is displayed on the display 112. In the illustrated embodiment, the image 502 comprises an avatar playing a representation of the guitar; para. [0058, 0072-0074]): and displaying, by the terminal apparatus, the reference image transmitted from the computer system, on the display of the terminal apparatus (i.e. The console 106 generates an image 502 that is representative of the user 102 playing the guitar 104. The image 502 is displayed on the display 112. In the illustrated embodiment, the image 502 comprises an avatar playing a representation of the guitar; para. [0058, 0072-0074]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Avitabile and Lee to include the feature of Villa. One would have been motivated to make this modification because the detected fingering could be presented visually to the user in an intuitive form showing how the fingering is performed on the instrument.
10. Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Avitabile in view of Lee, Villa, and further in view of Chen et al. (U.S. Patent Application Pub. No. US 20200327860 A1).
Claim 5: Avitabile, Lee, and Villa teach the method for processing information according to claim 4. Avitabile does not explicitly teach wherein: the display is worn on the head of the user, the showing the reference image includes showing on the display, by the computer system, a captured image representative of the virtual player and the virtual string instrument in a virtual space, and the captured image is taken by a virtual camera in a position and an orientation in the virtual space controlled based on a movement of the head of the user and is shown as the reference image.
Villa further teaches the virtual string instrument (i.e. the console deciphers what user manipulation of the musical instrument, such as the fingering, blowing, or other technique, is required to create the sound produced by the musical instrument. The console then applies the deciphered user manipulation of the musical instrument to an image, such as an avatar. In this scenario the avatar mimics the deciphered user manipulation of the musical instrument on a representation of the musical instrument; para. [0067, 0068]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Avitabile and Lee to include the feature of Villa. One would have been motivated to make this modification because the detected fingering could be presented visually to the user in an intuitive form showing how the fingering is performed on the instrument.
However, Chen teaches wherein: the display is worn on the head of the user (i.e. The head mounted display system includes a wearable body, a display unit and a processing unit; para. [0004]), the showing the reference image includes showing on the display, by the computer system, a captured image representative of the virtual player and the in a virtual space (i.e. Detailed description for the steps is provided as follows. In step S1, there can be a virtual camera for capturing images of the virtual environment, and the pose, the position and the orientation of the virtual camera and an eye of an avatar played by the user can be determined based on the pose, the position and the orientation of the head mounted display system 1. Therefore, when the user wears the wearable body 11, the display unit 12 can display the scene in the first-person perspective mode for the user firstly according to the images captured by the virtual camera; para. [0037]), and the captured image is taken by a virtual camera in a position and an orientation in the virtual space controlled based on a movement of the head of the user and is shown as the reference image (i.e. when the user experiences the virtual environment, the tracking unit 14 can track at least one of the position, the orientation and the pose of the head mounted display system 1 by tracking at least one of the position, the orientation and the pose of the wearable body 11. Furthermore, the activating command is generated when the tracking result of the tracking unit 14 meets the predetermined condition; para. [0038, 0039]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Avitabile, Lee, and Villa to include the feature of Chen. One would have been motivated to make this modification because positioning the virtual camera to capture the avatar and surrounding virtual environment and allowing the viewpoint to be adjusted so that the suer can better understand the avatar’s position or state.
11. Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Avitabile in view of Lee, and further in view of Dimitriadis et al. (U.S. Patent Application Pub. No. US 20230297777 A1)
Claim 8: Avitabile and Lee teach the method for processing information according to claim 1. Avitabile does not explicitly teach wherein: the input information includes identification information for any of a plurality of players, and the at least one generation model learns a relationship between: training input information for each of the plurality of players, the training input information including identification information for a corresponding player; and training fingering information indicative of fingering of the corresponding player.
Lee further teaches wherein: the at least one generation model learns a relationship between: training input information for each of the plurality of players (i.e. We report experiments measuring recognition accuracy on a dataset of several pieces of varying difficulty played by multiple pianists' and demonstrate that our approaches are able to detect pressed piano keys and the piano player's fingerings with an accuracy higher than baselines … We had five pianists record the pieces, including two professionals, two with medium skill, and one beginner; Section 1, 4.1), the training input information (i.e. To identify which finger is pressing each note on the keyboard, we frame the problem as object detection and employ the same architecture as above, except without octave classification or the audio stream (since audio signals provide no information about fingering); Section 3.2.1); and training fingering information indicative of fingering of the corresponding player (i.e. We then manually labeled each bounding box with the finger num-ber(s) that are currently pressing keys, and train the proposed network on this dataset … We collected 65 minutes of video also at 60 fps, and annotated them frame by frame for both notes and fingering using a combination of MIDI and manual labeling; Section 3.2.2, 4.1).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Avitabile to include the feature of Lee. One would have been motivated to make this modification because it improves the accuracy and robustness of determination of the user’s fingering, particularly where the image or audio information alone is ambiguous.
However, Dimitriadis teaches wherein: the input information includes identification information for any of a plurality of players (i.e. The training data set 54 includes multiple tuples of training data, each tuple including user identifier data 54A identifying a particular user, raw text data 54B from that particular user, and ground truth classification data 54C; para. [0021]), and the at least one generation model learns a relationship between: training input information for each of the plurality of players, the training input information including identification information for a corresponding player (i.e. The user-specific token sets 36 are inputted into the multi-user personalized NLP model 38 during training, along with the corresponding ground truth classifications 54C, the model is trained based on the training data set 54, and a trained multi-user NLP model 38A is outputted; para. [0022]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Avitabile and Lee to include the feature of Dimitriadis. One would have been motivated to make this modification because inserting user identifying information into the training input so that one model can learn personalized predictions for many users.
12. Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Avitabile in view of Lee, and further in view of Ljolje et al. (U.S. Patent Application Pub. No. US 20150248884 A1)
Claim 9: Avitabile and Lee teach the method for processing information according to claim 1. Avitabile does not explicitly teach wherein: the at least one generation model includes a plurality of generation models for different players, the generating fingering information includes generating, by the computer system, the fingering information by processing the acquired input information using any of the plurality of generation models, and each of the plurality of generation models is a model that learns a relationship between: the training input information; and the training fingering information indicative of fingering of a corresponding player from among the different players.
However, Lee further teaches the generating fingering information includes generating, by the computer system, the fingering information by processing the acquired input information using generation model (i.e. We present a first step towards developing an interactive piano tutoring system that can observe a student playing the piano and give feedback about hand movements and musical accuracy. In particular, we have two primary aims: 1) to determine which notes on a piano are being played at any moment in time, 2) to identify which finger is pressing each note. We introduce a novel two-stream convolutional neural network that takes video and audio inputs together for detecting pressed notes and finger presses. We formulate our two problems in terms of multi-task learning and extend a state-of-the-art object detection model to incorporate both audio and visual features. In addition, we introduce a novel finger identification solution based on pressed piano note information. We experimentally confirm that our approach is able to detect pressed piano keys and the piano player's fingers with a high accuracy; abs), and the generation model is a model that learns a relationship between: the training input information (i.e. We created new datasets of people playing the piano for training and testing our techniques. Figure 2 shows the pipeline that we used for generating our dataset given three different input files recorded while people played the piano: videos, MIDI files, and music scores. First, we extracted image frames and audio from the input video, and then applied the pre-trained hand detector on image frames to obtain bounding boxes of fingers. For the audio stream, we extracted MFCC features with 100 ms windows, and converted these into multi-channel images based on the first and second order derivatives, as described above. We extracted keypress information from MIDI synchronized with the video to create keypress ground truth labels; Section 4.1); and the training fingering information indicative of fingering of a corresponding player from among the different players (i.e. We report experiments measuring recognition accuracy on a dataset of several pieces of varying difficulty played by multiple pianists' and demonstrate that our approaches are able to detect pressed piano keys and the piano player's fingerings with an accuracy higher than baselines … We had five pianists record the pieces, including two professionals, two with medium skill, and one beginner; Section 1, 4.1).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Avitabile to include the feature of Lee. One would have been motivated to make this modification because it improves the accuracy and robustness of determination of the user’s fingering, particularly where the image or audio information alone is ambiguous.
However, Ljolje teaches wherein: the at least one generation model includes a plurality of generation models for different players (i.e. each SD model 304 a, 304 b, 304 c may correlate to a different user of the same or different communication device 310; para. [0034]), the generating information includes generating, by the computer system, the information by processing the acquired input information using any of the plurality of generation models (i.e. the selection module 302 may select 506 a quantity of SD models 304 a, 304 b, 304 c according to the user profile 314 and the usage log 316; para. [0044]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Avitabile and Lee to include the feature of Ljolje. One would have been motivated to make this modification because a generic model representing multiple users may not be well suited to particular individuals because individual users differ, thereby providing a personal model for each user and selecting among a plurality of user dependent models during processing.
13. Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Avitabile in view of Lee, and further in view of Beck (U.S. Patent Application Pub. No. US 20120017748 A1)
Claim 10: Avitabile and Lee teach the method for processing information according to claim 1. Avitabile further teaches wherein: the string instrument, a detector that detects playing by a player, and the fingering information is generated by using a result provided by the detector (i.e. a plurality of finger tip sensors 102A-D which can each be worn on a finger 104 of a user. Each of the finger tip sensors 102A-D are a different colour. As will be explained later, the tablet 100 is operable to run software which provides musical instrument tuition to the user. The tablet is able to receive information transmitted from the finger tip sensors 102A-102D when they are worn by the user as the user is playing a musical instrument, and the software is then able to provide feedback to the user about how they can improve their performance; para. [0033-0036]).
Avitabile does not explicitly teach the string instrument includes a detector.
However, Lee further teaches wherein: the instrument, a detector that detects playing by a player (i.e. special electronic instruments that can record the notes that a student plays, for example through MIDI … We recorded MIDI files while a person was playing the piano and then aligned them with music scores for annotating our dataset; Section 1, fig. 3), and the training fingering information is generated by using a result provided by the detector (i.e. We then manually labeled each bounding box with the finger num-ber(s) that are currently pressing keys, and train the proposed network on this dataset … We collected 65 minutes of video also at 60 fps, and annotated them frame by frame for both notes and fingering using a combination of MIDI and manual labeling; Section 3.2.2, 4.1).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the invention of Avitabile to include the feature of Lee. One would have been motivated to make this modification because it improves the accuracy and robustness of determination of the user’s fingering, particularly where the image or audio information alone is ambiguous.
However, Beck teaches wherein: the string instrument includes a detector that detects playing by a player (i.e. FIG. 2 illustrates exemplary components of instrument 104 for generating the signals. The user may interact with instrument 104 by using strings 202 extended over a fretboard 204. The user may press strings 202 on fretboard 204 by using fingers. Subsequently, a detector 206 detects the contact and generates digital signals; para. [0023]), and the fingering information is generated by using a result provided by the detector (i.e. The signals sent from transmitter 208 are received by a receiver 302 of processing device 106. Subsequently, a processor 304 analyzes the signals to generate musical notation. As discussed above, musical notation may be in the form of tablature. Further, the tablature may be a standard tablature of a hybrid tablature. The hybrid tablature may include the finger position and time or duration information of the contact; para. [0026]).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to modify the combination of Avitabile and Lee to include the feature of Beck. One would have been motivated to make this modification because a generic model representing multiple users may not be well suited to particular individuals because the techniques are desirable that can efficiently detect the positions of the fingers of a user. Moreover, techniques are desirable to use the finger positions to generate musical notation.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
Helms (Pub. No. US 11996070 B1), an isolated guitar string audio capture and visual string indication and chord-finger number overlay process and system that provides a visual way to show what strings are being pressed and played during a guitar instruction video is disclosed. The isolated guitar string audio capture and visual string indication and chord-finger number overlay process and system visibly shows the student what strings are being pressed and played in a given chord by the guitar instructor in the video, all from one camera angle.
It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331, 1332-33, 216 U.S.P.Q. 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 U.S.P.Q. 275, 277 (C.C.P.A. 1968)).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TAN TRAN whose telephone number is (303)297-4266. The examiner can normally be reached on Monday - Thursday - 8:00 am - 5:00 pm MT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Matt Ell can be reached on 571-270-3264. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TAN H TRAN/Primary Examiner, Art Unit 2141