DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Information Disclosure Statement
The information disclosure statement(s) (IDS(s)) submitted on 3/20/2026 is/are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement(s) is/are being considered by the examiner.
Response to Arguments
Applicant’s arguments, see pages 9-15, filed 6/9/2026, with respect to all claims, have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-4, 9, 11-12, and 14-17 are rejected under 35 U.S.C. 103 as unpatentable over Nakamura et al. (JP 2006152156 A, December 13, 2007), hereinafter Nakamura '156, in view of Kotsuji (JP 2019053170 A, April 4, 2019).
Regarding claim 1, Nakamura '156 teaches a computer-implemented information processing method (Nakamura '156 ¶0011: "a CPU 6 that controls the entire device, a ROM 7 that stores control programs executed by the CPU 6 and various table data, a RAM 8 that temporarily stores performance data, various input information and calculation results") comprising: determining, by a processor (Nakamura '156 ¶0018: "Next, the CPU 6 performs a predetermined image recognition process on the two image data stored in the image data storage area to recognize the position of each hand and finger of the performer in a predetermined direction (at least one direction, preferably multiple directions, from the up-and-down direction = vertical direction, left-and-right direction = the direction of the long side (pitch) of the keyboard 1a, and front-and-back direction = the direction of the short side (depth) of the keyboard 1a."), based on musical instrument information indicative of a musical instrument (Nakamura '156 ¶0015: "In this embodiment, the imaging device 5 is composed of two small cameras 5a and 5b, respectively, which are installed near both ends of the keyboard 1a, as shown in Figure 2."), a target part of a body of a first player, the first player playing the musical instrument indicated by the musical instrument information (Nakamura '156 ¶0011: "an imaging device 5 that mainly images the hands and fingers of the performer"; Nakamura '156 ¶0025: "While the keys are being pressed, the CPU 6 repeatedly performs the hand and/or finger position recognition process of the performer described in (2) above and the musical tone characteristic control parameter generation and supply process described in (3) above."; Nakamura '156 ¶0028: "In the hand-level position recognition process, as shown in Figure 4(a), the position of the left and right hands is recognized independently in at least one direction from the vertical, horizontal, and forward/backward directions."; Nakamura '156 ¶0029: "In the finger-level position recognition process, as shown in Figure 4(b), the position of each finger is recognized independently in at least one direction from the vertical, horizontal, and forward/backward directions.").
Nakamura '156 does not explicitly disclose acquiring, by the processor, based on the determined target part, image information indicative of imagery of the determined target part from image data from at least one camera.
However, Kotsuji teaches acquiring, by the processor (Kotsuji ¶0043: "After the setup is complete, when the musician begins playing, the control unit 4 instructs the camera unit 1 to start recording and video recording of the performance (step #14)."), based on the determined target part, image information indicative of imagery of the determined target part from image data from at least one camera (Kotsuji ¶0027: "Camera unit 1 is set up so that the instrument being played is within the shooting range. Camera crew 1 will film the parts of the instrument being played by the musician practicing. For example, if instrument 3 is a piano, camera unit 1 will film the hands of the person practicing the instrument. The image sensor 14 outputs an analog signal (image signal) at a constant interval. The camera module 15 generates performance video data 82 based on the image signal output by the image sensor 14.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer-implemented information processing method of Nakamura '156 by adding the image acquisition and audiovisual processing of Kotsuji to generate actual performance data that includes performance sound data and performance video data (Kotsuji ¶0009) to verify whether the hand movements, hand position, or finger dexterity are appropriate (Kotsuji ¶0007).
Regarding claim 2, Nakamura '156 (in view of Kotsuji) teaches a computer-implemented information processing method comprising the features of claim 1 as discussed above.
Kotsuji further teaches transmitting the acquired image information to an external apparatus (Kotsuji ¶0058: "The user may also view the composite data 9 using a portable communication device that the user owns." Kotsuji ¶0057: "The composite data 9 is video data in which the performance parts of the model player and the performance parts of the instrument learner are arranged side by side.").
Regarding claim 3, Nakamura '156 (in view of Kotsuji) teaches a computer-implemented information processing method comprising the features of claim 1 as discussed above.
Kotsuji further teaches that the determining the target part includes determining the target part based on the identified musical instrument information (Kotsuji ¶0017: "Furthermore, the musical score data 21 may specify the part of the musical piece that should be used for playing each note. For example, keyboard instruments are played with both hands. Examples of keyboard instruments include a piano and an electone. The left hand is generally used to play the lower notes. The right hand is generally used to play the lower notes. In the musical score data 21, a playing part may be defined for each note depending on the instrument 3.").
Regarding claim 4, Nakamura '156 (in view of Kotsuji) teaches a computer-implemented information processing method comprising the features of claim 3 as discussed above.
Kotsuji further teaches that the related information includes at least one of: information indicative of sounds emitted from the musical instrument; information indicative of imagery of the musical instrument; information indicative of a musical score for the musical instrument; or information indicative of a combination of the musical instrument and a lesson schedule for the musical instrument (Kotsuji ¶0017: "Furthermore, the musical score data 21 may specify the part of the musical piece that should be used for playing each note. For example, keyboard instruments are played with both hands. Examples of keyboard instruments include a piano and an electone. The left hand is generally used to play the lower notes. The right hand is generally used to play the lower notes. In the musical score data 21, a playing part may be defined for each note depending on the instrument 3.").
Regarding claim 9, Nakamura '156 (in view of Kotsuji) teaches a computer-implemented information processing method comprising the features of claim 1 as discussed above.
Kotsuji further teaches that determining the target part includes determining the target part based on the musical instrument information and sound information (Kotsuji ¶0026: "The microphone 11 collects the sound of the performance [and] the sound processor 12 converts the analog signal output by the microphone 11 into digital data (performance sound data 81)." The sound of the performance includes musical instrument information and sound information.), the sound information being indicative of sounds emitted from the musical instrument indicated by the musical instrument information (Kotsuji ¶0052: "The control unit 4 evaluates the performance according to the selected performance part… the control unit 4 compares the notes in the musical score data 21 that correspond to the direction of the selected hand with the performance sound data 81 and performs an evaluation").
Regarding claim 14, Nakamura '156 (in view of Kotsuji) teaches a computer-implemented information processing method comprising the features of claim 1 as discussed above.
Kotsuji further teaches that determining the target part includes determining the target part based on attention information indicative of an attention matter regarding playing the musical instrument (Kotsuji ¶0037: "The first setting screen 65a also has a third radio button RB3, a fourth radio button RB4, and a fifth radio button RB5. Any one of the third radio button RB3 to the fifth radio button RB5 is selected. When practicing with only the right hand, the third radio button RB3 is operated. When practicing with only the left hand, the fourth radio button RB4 is operated. When practicing with both hands, the fifth radio button RB5 is operated. The operation panel 6 accepts the selection of the body part to be used in practice.").
Regarding claim 15, Nakamura '156 (in view of Kotsuji) teaches a computer-implemented information processing method comprising the features of claim 1 as discussed above.
Kotsuji further teaches that determining the target part includes determining the target part based on player information regarding the first player (Kotsuji ¶0037: "The first setting screen 65a also has a third radio button RB3, a fourth radio button RB4, and a fifth radio button RB5. Any one of the third radio button RB3 to the fifth radio button RB5 is selected. When practicing with only the right hand, the third radio button RB3 is operated. When practicing with only the left hand, the fourth radio button RB4 is operated. When practicing with both hands, the fifth radio button RB5 is operated. The operation panel 6 accepts the selection of the body part to be used in practice.").
Regarding claim 17, Nakamura '156 (in view of Kotsuji) teaches a computer-implemented information processing method comprising the features of claim 1 as discussed above.
Kotsuji further teaches identifying the musical instrument information (Kotsuji ¶0033: "Next, the person practicing the instrument enters the instrument 3 to be practiced, the song to be practiced, and their name into the control panel 6. As a result, the control unit 4 recognizes the instrument 3 to be practiced, the song, and the person practicing the instrument (step #12)."), wherein the musical instrument information indicates a type of musical instrument played by the first player (Kotsuji ¶0033: "the person practicing the instrument enters the instrument 3 to be practiced"; Kotsuji ¶0038: "Note that the parts of the human body used for playing will vary depending on the instrument (instrument 3). For example, the Electone has foot pedals. Electone players operate the foot pedals with their left foot."), and wherein the determining the target part includes determining, as the target part, a part of a body that is associated with the type of musical instrument indicated by the musical instrument information (Kotsuji ¶0018: "In Figure 1, the piano is shown as an example of instrument 3. If instrument 3 is a piano, the part of the body used for playing is the hands (fingers). Camera unit 1 is positioned to capture images of the hand (fingers)."; Kotsuji ¶0038: "Thus, the number of radio buttons used to select the playing part may be changed depending on the instrument 3.").
Regarding claim 11, Nakamura '156 teaches a computer-implemented information processing method (Nakamura '156 ¶0011: "a CPU 6 that controls the entire device, a ROM 7 that stores control programs executed by the CPU 6 and various table data, a RAM 8 that temporarily stores performance data, various input information and calculation results") comprising: determining, by a processor (Nakamura '156 ¶0018: "Next, the CPU 6 performs a predetermined image recognition process on the two image data stored in the image data storage area to recognize the position of each hand and finger of the performer in a predetermined direction (at least one direction, preferably multiple directions, from the up-and-down direction = vertical direction, left-and-right direction = the direction of the long side (pitch) of the keyboard 1a, and front-and-back direction = the direction of the short side (depth) of the keyboard 1a."), based on sound information indicative of sounds emitted from a musical instrument (Nakamura '156 ¶0018: "When the performer starts pressing the keyboard 1a, the CPU 6 first generates key-on event data corresponding to the key press and supplies it to the sound source circuit 14. In response, the sound source circuit 14 starts to produce a musical tone corresponding to the supplied key-on event data."), a target part of a body of a first player, the first player playing the musical instrument (Nakamura '156 ¶0011: "an imaging device 5 that mainly images the hands and fingers of the performer"; Nakamura '156 ¶0025: "While the keys are being pressed, the CPU 6 repeatedly performs the hand and/or finger position recognition process of the performer described in (2) above and the musical tone characteristic control parameter generation and supply process described in (3) above."; Nakamura '156 ¶0028: "In the hand-level position recognition process, as shown in Figure 4(a), the position of the left and right hands is recognized independently in at least one direction from the vertical, horizontal, and forward/backward directions."; Nakamura '156 ¶0029: "In the finger-level position recognition process, as shown in Figure 4(b), the position of each finger is recognized independently in at least one direction from the vertical, horizontal, and forward/backward directions.").
Nakamura '156 does not explicitly disclose acquiring, by the processor, based on the determined target part, image information indicative of imagery of the determined target part from image data from at least one camera.
However, Kotsuji teaches acquiring, by the processor (Kotsuji ¶0043: "After the setup is complete, when the musician begins playing, the control unit 4 instructs the camera unit 1 to start recording and video recording of the performance (step #14)."), based on the determined target part, image information indicative of imagery of the determined target part from image data from at least one camera (Kotsuji ¶0027: "Camera unit 1 is set up so that the instrument being played is within the shooting range. Camera crew 1 will film the parts of the instrument being played by the musician practicing. For example, if instrument 3 is a piano, camera unit 1 will film the hands of the person practicing the instrument. The image sensor 14 outputs an analog signal (image signal) at a constant interval. The camera module 15 generates performance video data 82 based on the image signal output by the image sensor 14.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer-implemented information processing method of Nakamura '156 by adding the image acquisition and audiovisual processing of Kotsuji to generate actual performance data that includes performance sound data and performance video data (Kotsuji ¶0009) to verify whether the hand movements, hand position, or finger dexterity are appropriate (Kotsuji ¶0007).
Regarding claim 12, Nakamura '156 (in view of Kotsuji) teaches a computer-implemented information processing method comprising the features of claim 11.
Kotsuji further teaches that determining the target part includes determining the target part based on a relationship between the sound information and musical-score information indicative of a musical score (Kotsuji ¶0017: "Furthermore, the musical score data 21 may specify the part of the musical piece that should be used for playing each note. For example, keyboard instruments are played with both hands. Examples of keyboard instruments include a piano and an electone. The left hand is generally used to play the lower notes. The right hand is generally used to play the lower notes. In the musical score data 21, a playing part may be defined for each note depending on the instrument 3.").
Regarding claim 16, Nakamura '156 teaches an information processing system comprising: at least one memory configured to store instructions (Nakamura '156 ¶0011: "a CPU 6 that controls the entire device, a ROM 7 that stores control programs executed by the CPU 6 and various table data, a RAM 8 that temporarily stores performance data, various input information and calculation results"); and at least one processor configured to execute the instructions (Nakamura '156 ¶0018: "Next, the CPU 6 performs a predetermined image recognition process on the two image data stored in the image data storage area to recognize the position of each hand and finger of the performer in a predetermined direction (at least one direction, preferably multiple directions, from the up-and-down direction = vertical direction, left-and-right direction = the direction of the long side (pitch) of the keyboard 1a, and front-and-back direction = the direction of the short side (depth) of the keyboard 1a.") to: determine, based on musical instrument information indicative of a musical instrument (Nakamura '156 ¶0015: "In this embodiment, the imaging device 5 is composed of two small cameras 5a and 5b, respectively, which are installed near both ends of the keyboard 1a, as shown in Figure 2."), a target part of a body of a first player, the first player playing the musical instrument indicated by the musical instrument information (Nakamura '156 ¶0011: "an imaging device 5 that mainly images the hands and fingers of the performer"; Nakamura '156 ¶0025: "While the keys are being pressed, the CPU 6 repeatedly performs the hand and/or finger position recognition process of the performer described in (2) above and the musical tone characteristic control parameter generation and supply process described in (3) above."; Nakamura '156 ¶0028: "In the hand-level position recognition process, as shown in Figure 4(a), the position of the left and right hands is recognized independently in at least one direction from the vertical, horizontal, and forward/backward directions."; Nakamura '156 ¶0029: "In the finger-level position recognition process, as shown in Figure 4(b), the position of each finger is recognized independently in at least one direction from the vertical, horizontal, and forward/backward directions.").
Nakamura '156 does not explicitly disclose: acquire, based on the determined target part, image information indicative of imagery of the determined target part from image data from at least one camera.
However, Kotsuji teaches: acquire (Kotsuji ¶0043: "After the setup is complete, when the musician begins playing, the control unit 4 instructs the camera unit 1 to start recording and video recording of the performance (step #14)."), based on the determined target part, image information indicative of imagery of the determined target part from image data from at least one camera (Kotsuji ¶0027: "Camera unit 1 is set up so that the instrument being played is within the shooting range. Camera crew 1 will film the parts of the instrument being played by the musician practicing. For example, if instrument 3 is a piano, camera unit 1 will film the hands of the person practicing the instrument. The image sensor 14 outputs an analog signal (image signal) at a constant interval. The camera module 15 generates performance video data 82 based on the image signal output by the image sensor 14.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer-implemented information processing system of Nakamura '156 by adding the image acquisition and audiovisual processing of Kotsuji to generate actual performance data that includes performance sound data and performance video data (Kotsuji ¶0009) to verify whether the hand movements, hand position, or finger dexterity are appropriate (Kotsuji ¶0007).
Claims 5-6, 10, and 13 are rejected under 35 U.S.C. 103 as unpatentable over Nakamura '156 in view of Kotsuji and further in view of Dorfer et al. ("Towards Score Following in Sheet Music Images," https://arxiv.org/pdf/1612.05050, December 15, 2016, retrieved February 20, 2026), hereinafter Dorfer.
Regarding claim 5, Nakamura '156 (in view of Kotsuji) teaches a computer-implemented information processing method comprising the features of claim 3 as discussed above.
Nakamura '156 (in view of Kotsuji) does not explicitly disclose that identifying the musical instrument information includes: inputting the related information into a trained model, the trained model having been trained to learn a relationship between training-related information and training-musical-instrument information, the training-related information being related to the musical instrument, and the training-musical-instrument information being indicative of a musical instrument specified from the training-related information; and identifying, as the musical instrument information, information output from the trained model in response to the related information.
However, Dorfer teaches or suggests that identifying the musical instrument information includes: inputting the related information into a trained model (Dorfer abstract: "It consists of an end-to-end multi-modal convolutional neural network that takes as input images of sheet music and spectrograms of the respective audio snippets"), the trained model having been trained to learn a relationship between training-related information and training-musical-instrument information (Dorfer § 2.1: "The model takes two different input modalities at the same time: images of scores, and short excerpts from spectrograms of audio renditions of the score"), the training-related information being related to the musical instrument (Dorfer § 3.2: "We select the first track of the midi files (right hand, piano) and render it as sheet music using Lilypond."), and the training-musical-instrument information being indicative of a musical instrument specified from the training-related information (Dorfer § 3.2: "We synthesize the midi-tracks to flac-audio using Fluidsynth 4 and a Steinway piano sound font."); and identifying, as the musical instrument information, information output from the trained model in response to the related information (Dorfer § 4.2: "As a final point, we report on first attempts at working with 'real' music. For this purpose one of the authors played the right hand part of a simple piece (Minuet in G Major by Johann Sebastian Bach, BWV Anhang 114) – which, of course, was not part of the training data – on a Yamaha AvantGrand N2 hybrid piano").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer-implemented information processing method of Nakamura '156 (as modified by Kotsuji) by adding the trained model of Dorfer to precisely link a performance to its respective sheet music (Dorfer § 1).
Regarding claim 6, Nakamura '156 (in view of Kotsuji and further in view of Dorfer) teaches a computer-implemented information processing method comprising the features of claim 5 as discussed above.
Dorfer further teaches or suggests that the related information and the training-related information each indicates sounds emitted from the musical instrument (Dorfer § 2.1: "The model takes two different input modalities at the same time: images of scores, and short excerpts from spectrograms of audio renditions of the score"); and the training-musical-instrument information indicates, as the musical instrument specified from the training-related information, a musical instrument that emits the sounds indicated by the training-related information (Dorfer § 3.2: "We synthesize the midi-tracks to flac-audio using Fluidsynth 4 and a Steinway piano sound font.").
Regarding claim 10, Nakamura '156 (in view of Kotsuji) teaches a computer-implemented information processing method comprising the features of claim 9 as discussed above.
Nakamura '156 (in view of Kotsuji) does not explicitly disclose that determining the target part includes: inputting input information into a trained model, the input information including the musical instrument information and the sound information, the trained model having been trained to learn a relationship between training-input information and training-output information, the training-input information including training-musical-instrument information and training-sound information, the training-musical-instrument information being indicative of the musical instrument, the training-sound information being indicative of sounds emitted from the musical instrument indicated by the training-musical-instrument information, the training-output information being indicative of a target part of a body of a second player, the second player playing the musical instrument indicated by the training-musical-instrument information, and the musical instrument indicated by the training-musical-instrument information emitting the sounds indicated by the training-sound information; and determining the target part based on output information output from the trained model in response to the input information.
However, Dorfer teaches or suggests that that determining the target part includes: inputting input information into a trained model (Dorfer abstract: "It consists of an end-to-end multi-modal convolutional neural network that takes as input images of sheet music and spectrograms of the respective audio snippets"), the input information including the musical instrument information (Dorfer § 3.2: "We select the first track of the midi files (right hand, piano) and render it as sheet music using Lilypond.") and the sound information (Dorfer § 3.2: "We synthesize the midi-tracks to flac-audio using Fluidsynth 4 and a Steinway piano sound font."), the trained model having been trained to learn a relationship between training-input information and training-output information (Dorfer abstract: "It learns to predict, for a given unseen audio snippet (covering approximately one bar of music), the corresponding position in the respective score line."), the training-input information including training-musical-instrument information and training-sound information (Dorfer § 2.1: "The model takes two different input modalities at the same time: images of scores, and short excerpts from spectrograms of audio renditions of the score"), the training-musical-instrument information being indicative of the musical instrument (Dorfer § 3.2: "We select the first track of the midi files (right hand, piano) and render it as sheet music using Lilypond."), the training-sound information being indicative of sounds emitted from the musical instrument indicated by the training-musical-instrument information (Dorfer § 3.2: "We synthesize the midi-tracks to flac-audio using Fluidsynth 4 and a Steinway piano sound font."), the training-output information being indicative of a target part of a body of a second player (Dorfer § 3.2: "We select the first track of the midi files (right hand, piano) and render it as sheet music using Lilypond."), the second player playing the musical instrument indicated by the training-musical-instrument information (Dorfer § 4.2: "As a final point, we report on first attempts at working with 'real' music. For this purpose one of the authors played the right hand part of a simple piece"), and the musical instrument indicated by the training-musical-instrument information emitting the sounds indicated by the training-sound information (Dorfer § 4.2: "As a final point, we report on first attempts at working with 'real' music. For this purpose one of the authors played the right hand part of a simple piece (Minuet in G Major by Johann Sebastian Bach, BWV Anhang 114) – which, of course, was not part of the training data – on a Yamaha AvantGrand N2 hybrid piano"); and determining the target part based on output information output from the trained model in response to the input information (Dorfer § 4.2: "As a final point, we report on first attempts at working with 'real' music. For this purpose one of the authors played the right hand part of a simple piece (Minuet in G Major by Johann Sebastian Bach, BWV Anhang 114) – which, of course, was not part of the training data – on a Yamaha AvantGrand N2 hybrid piano and recorded it using a single microphone. In this application scenario we predict the corresponding sheet locations not only at times of onsets but for a continuous audio stream (subsequent spectrogram excerpts).").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer-implemented information processing method of Nakamura '156 (as modified by Kotsuji) by adding the trained model of Dorfer to precisely link a performance to its respective sheet music (Dorfer § 1).
Regarding claim 13, Nakamura '156 (in view of Kotsuji) teaches a computer-implemented information processing method comprising the features of claim 11.
Kotsuji further suggests the second player playing the musical instrument in accordance with the musical score indicated by the training-musical-score information (Kotsuji ¶0009: "the playing parts of a musical instrument practitioner playing the piece of music corresponding to the musical score data."), and the musical instrument being capable of emitting the sounds indicated by the training-sound information (Kotsuji ¶0080: "The synthesized data 9 may include musical sounds. When the composite data 9 is played back, the control unit 4 may simultaneously play back the performance of the musical instrument learner from the speaker 64").
Nakamura '156 (in view of Kotsuji) does not explicitly disclose that determining the target part includes: inputting input information into a trained model, the input information including the sound information and musical-score information, the musical-score information being indicative of a musical score, the trained model having been trained to learn a relationship between training-input information and training-output information, the training-input information including training-sound information and training-musical-score information, the training-sound information being indicative of sounds emitted from the musical instrument, the training-musical-score information being indicative of a musical score, the training-output information being indicative of a target part of a body of a second player; and determining the target part based on output information output from the trained model in response to the input information.
However, Dorfer teaches or suggests that determining the target part includes: inputting input information into a trained model, the input information including the sound information and musical-score information (Dorfer abstract: "It consists of an end-to-end multi-modal convolutional neural network that takes as input images of sheet music and spectrograms of the respective audio snippets"), the musical-score information being indicative of a musical score (Dorfer abstract: "takes as input images of sheet music "), the trained model having been trained to learn a relationship between training-input information and training-output information (Dorfer abstract: "It learns to predict, for a given unseen audio snippet (covering approximately one bar of music), the corresponding position in the respective score line."), the training-input information including training-sound information and training-musical-score information (Dorfer § 2.1: "The model takes two different input modalities at the same time: images of scores, and short excerpts from spectrograms of audio renditions of the score"), the training-sound information being indicative of sounds emitted from the musical instrument (Dorfer § 3.2: "We synthesize the midi-tracks to flac-audio using Fluidsynth 4 and a Steinway piano sound font."), the training-musical-score information being indicative of a musical score (Dorfer § 3.2: "We select the first track of the midi files (right hand, piano) and render it as sheet music using Lilypond."), the training-output information being indicative of a target part of a body of a second player (Dorfer § 3.2: "We select the first track of the midi files (right hand, piano) and render it as sheet music using Lilypond."); and determining the target part based on output information output from the trained model in response to the input information (Dorfer § 4.2: "As a final point, we report on first attempts at working with 'real' music. For this purpose one of the authors played the right hand part of a simple piece (Minuet in G Major by Johann Sebastian Bach, BWV Anhang 114) – which, of course, was not part of the training data – on a Yamaha AvantGrand N2 hybrid piano and recorded it using a single microphone. In this application scenario we predict the corresponding sheet locations not only at times of onsets but for a continuous audio stream (subsequent spectrogram excerpts).").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer-implemented information processing method of Nakamura '156 (as modified by Kotsuji) by adding the trained model of Dorfer to precisely link a performance to its respective sheet music (Dorfer § 1).
Claim 7 is rejected under 35 U.S.C. 103 as unpatentable over Nakamura '156 in view of Kotsuji and further in view of Dorfer and Arandjelovi et al. "Objects that Sound," https://arxiv.org/pdf/1712.06651, July 25, 2018, Retrieved February 20, 2026), hereinafter Arandjelovi.
Regarding claim 7, Nakamura '156 (in view of Kotsuji and further in view of Dorfer) teaches a computer-implemented information processing method comprising the features of claim 5 as discussed above.
Nakamura '156 (in view of Kotsuji and further in view of Dorfer) does not explicitly disclose that the related information and the training-related information each indicates imagery of the musical instrument, and the training-musical-instrument information indicates, as the musical instrument specified from the training-related information, a musical instrument represented by the imagery indicated by the training-related information.
However, Arandjelovi teaches or suggests that the related information and the training-related information each indicates imagery of the musical instrument (Arandjelovi § 3.1: "The architectures are trained on the AudioSet-Instruments train-val set, and evaluated on the AudioSet-Instruments test set described in Section 2"), and the training-musical-instrument information indicates, as the musical instrument specified from the training-related information, a musical instrument represented by the imagery indicated by the training-related information (§ 4.1: "The ability of the network to localize the object(s) that sound is demonstrated in Figure 5. It is able to detect a wide range of objects in different viewpoints and scales, and under challenging imaging conditions. A more detailed discussion including the analysis of some failure cases is available in the figure caption. As expected from an unsupervised method, it is not necessarily the case that it detects the entire object but can focus only on specific discriminative parts such as the interface between the hands and the piano keyboard. ").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer-implemented information processing method of Nakamura '156 (as modified by Kotsuji and Dorfer) by adding the trained model of Arandjelovi to focus on specific discriminative parts of an instrument such as the interface between the hands and the piano keyboard (Arandjelovi § 4.1).
Claim 8 is rejected under 35 U.S.C. 103 as unpatentable over Nakamura '156 in view of Kotsuji and further in view of Takijiri et al. (JP 2014167576 A, September 11, 2014), hereinafter Takijiri.
Regarding claim 8, Nakamura '156 (in view of Kotsuji) teaches a computer-implemented information processing method comprising the features of claim 3 as discussed above.
Nakamura '156 (in view of Kotsuji) does not explicitly disclose that identifying the musical instrument information includes identifying, as the musical instrument information, reference-musical-instrument information associated with the related information by referring to a table indicative of associations between reference-related information related to the musical instrument and the reference-musical-instrument information indicative of the musical instrument.
However, Takijiri teaches or suggests that identifying the musical instrument information includes identifying, as the musical instrument information, reference-musical-instrument information associated with the related information by referring to a table indicative of associations between reference-related information related to the musical instrument and the reference-musical-instrument information indicative of the musical instrument (Takijiri ¶0030: "Here, the image recognition process performed when setting the performance part in FIG. 3 (a) will be described with reference to FIGS. The control device 20 has the instrument image table shown in FIG. 4 stored in the ROM 47, the CD-ROM 53, or the like. In this instrument image table, exterior images of various instruments that can be used in minus-one performances are stored in association with the names of the respective instruments... In this embodiment, the karaoke device 10 performs image recognition processing to identify the type of instrument by comparing the external image of the instrument captured by the camera 25 with which external image in the instrument image table it is closest to.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer-implemented information processing method of Nakamura '156 (as modified by Kotsuji) by adding the table of Takijiri to identify the type of instrument captured by the camera (Takajiri ¶0030).
Claim 18 is rejected under 35 U.S.C. 103 as unpatentable over Nakamura '156 in view of Kotsuji and further in view of Nishimura (JP 2016038529 A, March 22, 2016) and Nakamura et al. ("Statistical Learning and Estimation of Piano Fingering," January 1, 2020, retrieved September 20, 2026 from https://arxiv.org/pdf/1904.10237), hereinafter Nakamura paper.
Regarding claim 18, Nakamura '156 (in view of Kotsuji) teaches a computer-implemented information processing method comprising the features of claim 1 as discussed above.
Kotsuji further teaches changing, by the processor, the target part to the updated target part (Kotsuji ¶0052: "At this point, the control panel 6 accepts the selection of the part to be played. The control unit 4 evaluates the performance according to the selected performance section."); and acquiring, by the processor, updated image information indicative of imagery of the updated target part from the image data from the at least one camera (Kotsuji ¶0043: "After the setup is complete, when the musician begins playing, the control unit 4 instructs the camera unit 1 to start recording and video recording of the performance (step #14)."; Kotsuji ¶0027: "Camera unit 1 is set up so that the instrument being played is within the shooting range. Camera crew 1 will film the parts of the instrument being played by the musician practicing.").
Nakamura '156 (in view of Kotsuji) does not explicitly disclose determining, by the processor, an updated target part for an upcoming portion of a piece of music played by the first player based on musical-score data, wherein the determining the updated target part includes inputting information indicative of a current portion of the piece of music into a trained model that has been trained to output target part information for the upcoming portion of the piece of music.
However, Nishimura teaches determining, by the processor, an updated target part for an upcoming portion of a piece of music played by the first player based on musical-score data (Nishimura ¶0007: "a key position acquisition means for sequentially acquiring the position of the key on the keyboard to be played for each note that constitutes a musical piece"; Nishimura ¶0025: "An on command indicating a key press includes a key number specifying the key to be played and a finger number specifying the finger to press the key."; Nishimura ¶0031: "In the guide process, when the command timing (the timing when the timer value of register T becomes '0' or less) is reached, the guide data GD, which is the 'command,' is read. If it is an 'on command', the key number NOTE and finger number FINGER are extracted, and the on flag ONF is set to '1.'").
Furthermore, Nakamura paper teaches that the determining the updated target part includes inputting information indicative of a current portion of the piece of music into a trained model that has been trained to output target part information for the upcoming portion of the piece of music (Nakamura paper § 1: "In this paper, we present a newly released dataset of piano fingering and propose fingering estimation methods based on hidden Markov models (HMMs) and DNNs."; Nakamura paper § 5.1: "The input is a piano performance represented as a MIDI-like signal (pn,tn,¯tn)N n=1 where each musical note n is described by its pitch pn, onset time tn, and offset time ¯tn (N denotes the number of notes)… The output is a list of finger numbers (fn)N n=1 corresponding to each note n.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer-implemented information processing method of Nakamura '156 (as modified by Kotsuji) by adding the updated target part of Nishimura and the trained model of Nakamura paper to define the part to be played for each note (Kotsuji ¶0017) and to compare the musical score data and the performance during the performance (Kotsuji ¶0067).
Claim 19 is rejected under 35 U.S.C. 103 as unpatentable over Nakamura '156 in view of Kotsuji and further in view of Sengupta et al. (US 6359647 B1, March 19, 2002), hereinafter Sengupta.
Regarding claim 19, Nakamura '156 (in view of Kotsuji) teaches a computer-implemented information processing method comprising the features of claim 1 as discussed above.
Nakamura '156 (in view of Kotsuji) does not explicitly disclose that the acquiring includes determining, by the processor, at least one usable camera from among a plurality of cameras based on the determined target part, and acquiring the image information from image data generated by the at least one usable camera.
However, Sengupta teaches that the acquiring includes determining, by the processor, at least one usable camera from among a plurality of cameras based on the determined target part (Sengupta col. 9, lines 53-66: "Given the target point P, in the physical domain coordinate system, at 650, the camera handoff system in accordance with this invention determines which cameras contain this target point in the loop 652-658. Because each camera's potential field of view is represented as vertices in the physical domain coordinate system, the process merely comprises a determination of whether point P lies within the polygon or polyhedron (CamPoly) associated with each camera. If the number of cameras is large, this search process can be optimized, as discussed above. The search process employed would replace this exhaustive search loop at 652-658. The system thereafter selects one of the cameras containing the target point P in its potential field of view, at 660."), and acquiring the image information from image data generated by the at least one usable camera (Sengupta col. 4, lines 35-42: "In accordance with this invention, based upon the determined location of the person and the determined field of view of each camera, the controller 130 selects camera 106 when the person enters camera 106's potential field of view. In a preferred embodiment that includes a figure tracking system 144, the figure tracking techniques will subsequently be applied to continue to track the figure in the image from camera 106.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer-implemented information processing method of Nakamura '156 (as modified by Kotsuji) by adding the plurality of cameras of Sengupta to provide for the automation of a multiple camera system, so as to provide for a multi-camera figure tracking capability to allow for the near continuous display of a figure as the figure moves about throughout the multiple cameras' potential fields of view (Sengupta col. 1, line 65 – col. 2, line 3).
Claim 20 is rejected under 35 U.S.C. 103 as unpatentable over Nakamura '156 in view of Kotsuji and further in view of Cao et al. ("OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields," April 14, 2017, retrieved September 20, 2026 from https://arxiv.org/pdf/1611.08050), hereinafter Cao.
Regarding claim 20, Nakamura '156 (in view of Kotsuji) teaches a computer-implemented information processing method comprising the features of claim 1 as discussed above.
Nakamura '156 (in view of Kotsuji) does not explicitly disclose that the acquiring includes extracting, by the processor, a portion of the image data corresponding to the determined target part from an image of a body of the first player captured by the at least one camera by using an image recognition technique.
However, Cao teaches that the acquiring includes extracting, by the processor, a portion of the image data corresponding to the determined target part from an image of a body of the first player captured by the at least one camera by using an image recognition technique (Cao abstract: "We present an approach to efficiently detect the 2D pose of multiple people in an image."; Cao § 2: "First, a feed-forward network simultaneously predicts a set of 2D confidence maps S of body part locations (Fig. 2b) and a set of 2D vector fields L of part affinities, which encode the degree of association between parts (Fig. 2c)."; Cao § 2.2: "Each confidence map is a 2D representation of the belief that a particular body part occurs at each pixel location… At test time, we predict confidence maps (as shown in the first row of Fig. 4), and obtain body part candidates by performing non-maximum suppression.").
It would have been prima facie obvious to one of ordinary skill in the art prior to the effective filing date of the claimed invention to have modified the computer-implemented information processing method of Nakamura '156 (as modified by Kotsuji) by adding image recognition technique of Cao to generate actual performance data that includes performance sound data and performance video data (Kotsuji ¶0009) to verify whether the hand movements, hand position, or finger dexterity are appropriate (Kotsuji ¶0007).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHILIP SCOLES whose telephone number is (703)756-1831. The examiner can normally be reached Monday-Friday 8:30-4:30 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Dedei Hammond can be reached on 571-270-7938. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PHILIP G SCOLES/
Examiner, Art Unit 2837
/DEDEI K HAMMOND/Supervisory Patent Examiner, Art Unit 2837