Prosecution Insights
Last updated: August 18, 2026
Application No. 19/024,265

VOICE FEEDBACK FOR USER INTERFACE OF MEDIA PLAYBACK DEVICE

Non-Final OA §102§103
Filed
Jan 16, 2025
Priority
Nov 01, 2018 — continuation of 11/043,216 +1 more
Examiner
MARLOW, ALEXANDER G
Art Unit
Tech Center
Assignee
Spotify AB
OA Round
1 (Non-Final)
79%
Grant Probability
Favorable
1-2
OA Rounds
1y 1m
Est. Remaining
97%
With Interview

Examiner Intelligence

Grants 79% — above average
79%
Career Allowance Rate
66 granted / 84 resolved
+18.6% vs TC avg
Strong +18% interview lift
Without
With
+18.1%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
7 currently pending
Career history
90
Total Applications
across all art units

Statute-Specific Performance

§101
16.9%
-23.1% vs TC avg
§103
50.3%
+10.3% vs TC avg
§102
16.2%
-23.8% vs TC avg
§112
10.2%
-29.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 84 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Introduction This office action is in response to communications filed 01/16/2025. Claims 2-21 are pending and likewise have been examined. Information Disclosure Statement The information disclosure statement (IDS) submitted on 04/28/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Objections Claim 7 is objected to because of the following informalities: Claim 7 recites “the group” on Line 1-2. The antecedent basis is improper. While the claim is clear on intending to establish a new group, it would be proper if the claim recited “a group”, as the group has not been established before. Appropriate correction is required. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13. The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer. Claims 2-21 rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-15 of U.S. Patent No. 11043216. Although the claims at issue are not identical, they are not patentably distinct from each other. See the table below for mapping of similar limitations. Instant Application U.S. Patent No. 11043216 Claim 2: A method of providing voice feedback, comprising: Claim 1: A method of providing voice feedback to a listener as part of a user interface of a media playback system, the method comprising: storing multiple different voice feedback recordings in at least one computer-readable storage device; storing multiple different voice feedback recordings in at least one computer-readable storage device receiving a listener command corresponding to a musical selection; receiving, with the media playback system, a listener command corresponding to a musical selection; determining, with a processing device, an identifying musical characteristic of the musical selection; determining, with a processing device of the media playback system, an identifying musical characteristic of the musical selection; selecting a first voice feedback recording from the multiple different voice feedback recordings, using the processing device, wherein the first voice feedback recording corresponds to the identifying musical characteristic; selecting a first voice feedback recording from the multiple different voice feedback recordings, using the processing device, wherein the first voice feedback recording corresponds to the identifying musical characteristic; and causing playback of the first voice feedback recording and the musical selection via a media playback system. and playing the first voice feedback recording to the listener on-beat with the musical selection via the media playback system. Claim 3: The method of claim 2, wherein each of the multiple different voice feedback recordings corresponds to a different style of music. Claim 2: The method of claim 1, wherein each of the multiple different voice feedback recordings of the different voice artists corresponds to a different style of music, Claim 4: The method of claim 2, wherein the identifying musical characteristic comprises a particular style of music selected from a predefined list of different styles of music. Claim 2: and wherein the identifying musical characteristic comprises a particular style of music selected from a predefined list of different styles of music. Claim 5: The method of claim 2, further comprising, before the storing step: receiving a first voice recording; and generating a first set of multiple voice recordings from the first voice recording using at least one of machine learning or natural language generation. Claim 3: The method of claim 2, further comprising, before the storing step: receiving a first voice recording from a first voice artist; and generating a first set of multiple voice recordings from the first voice recording, using at least one of machine learning or natural language generation. Claim 6: The method of claim 5, wherein the first set of multiple voice recordings comprises at least one of different tempos, different words, different pitches or different speaking styles. Claim 4: The method of claim 3, wherein the first set of multiple recordings comprises at least one of different tempos, different words, different pitches and different speaking styles of recordings of the first voice artist. Claim 7: The method of claim 2, wherein the musical selection is selected from the group consisting of a piece of music, an album, an artist, a style of music, a playlist, a shelf of music and a card of music. Claim 5: The method of claim 1, wherein the musical selection is selected from the group consisting of a piece of music, an album, an artist, a style of music, a playlist, a shelf of music and a card of music. Claim 8: The method of claim 2, wherein storing the multiple different voice feedback recordings comprises storing different tempo recordings. Claim 6: The method of claim 1, wherein storing the multiple different voice feedback recordings comprises storing different tempo recordings for each voice artist. Claim 9: The method of claim 2, wherein receiving the listener command comprises receiving at least one of a shelf selection or a card selection. Claim 7: The method of claim 1, wherein receiving the listener command comprises receiving at least one of a shelf selection or a card selection. Claim 10: The method of claim 2, further comprising: creating a voice beat grid for the first voice feedback recording; and creating a music beat grid for the musical selection. Claim 8: The method of claim 1, further comprising: creating a voice beat grid for the first voice feedback recording; and creating a music beat grid for the musical selection. Claim 11: The method of claim 2, wherein the first voice feedback recording is played at least partially before the musical selection is played by the media playback system. Claim 9: The method of claim 1, wherein the first voice feedback recording is played at least partially before the musical selection is played by the media playback system. Claim 12: The method of claim 11, wherein a portion of the first voice feedback recording is played at a same time that a beginning portion of the musical selection is played. Claim 10: The method of claim 9, wherein a portion of the first voice feedback recording is played at a same time that a beginning portion of the musical selection is played. Claim 13: The method of claim 12, wherein at least the portion of the first voice feedback recording is played on-beat with the musical selection. Claim 1: The method of claim 10, wherein at least the portion of the first voice feedback recording is played on-beat with the musical selection. Claim 14: The method of claim 2, wherein the first voice feedback recording is played at least partially after the musical selection is played by the media playback system. Claim 12: The method of claim 1, wherein the first voice feedback recording is played at least partially after the musical selection is played by the media playback system. Claim 15: The method of claim 2, wherein the multiple different voice feedback recordings comprise multiple introductions of multiple possible musical selections. Claim 13: The method of claim 1, wherein the multiple different voice feedback recordings comprise multiple introductions of multiple possible musical selections. Claim 16: The method of claim 2, further comprising customizing at least the first voice feedback recording to address a listener by name. Claim 14: The method of claim 1, further comprising customizing at least the first voice feedback recording to address the listener by name. Claim 17: A non-transitory computer readable medium for use on a computer system containing computer-executable programming instructions, the instructions including instructions for: Claim 15: A non-transitory computer readable medium for use on a computer system containing computer-executable programming instructions for providing voice feedback to a listener as part of a user interface of a media playback system, the instructions being executable by the computer system to: storing multiple different voice feedback recordings in at least one computer-readable storage device; store multiple different voice feedback recordings in at least one computer-readable storage device receiving a listener command corresponding to a musical selection; determining, with a processing device, an identifying musical characteristic of the musical selection; receive, with the media playback system, a listener command corresponding to a musical selection; determine, with a processing device of the media playback system, an identifying musical characteristic of the musical selection selecting a first voice feedback recording from the multiple different voice feedback recordings, using the processing device, wherein the first voice feedback recording corresponds to the identifying musical characteristic; select a first voice feedback recording from the multiple different voice feedback recordings, using the processing device, wherein the first voice feedback recording corresponds to the identifying musical characteristic; and causing playback of the first voice feedback recording and the musical selection via a media playback system. and play the first voice feedback recording to the listener on-beat with the musical selection via the media playback system. Claim 18: A system, comprising: one or more processors: memory storing instructions for execution by the one or more processors, including instructions for: Claim 1: A method of providing voice feedback to a listener as part of a user interface of a media playback system, the method comprising: See Note #1 below storing multiple different voice feedback recordings in at least one computer-readable storage device; storing multiple different voice feedback recordings in at least one computer-readable storage device receiving a listener command corresponding to a musical selection; receiving, with the media playback system, a listener command corresponding to a musical selection; determining, with a processing device, an identifying musical characteristic of the musical selection; determining, with a processing device of the media playback system, an identifying musical characteristic of the musical selection; selecting a first voice feedback recording from the multiple different voice feedback recordings, using the processing device, wherein the first voice feedback recording corresponds to the identifying musical characteristic; selecting a first voice feedback recording from the multiple different voice feedback recordings, using the processing device, wherein the first voice feedback recording corresponds to the identifying musical characteristic; and causing playback of the first voice feedback recording and the musical selection via a media playback system. and playing the first voice feedback recording to the listener on-beat with the musical selection via the media playback system. Claim 19: The system of claim 18, wherein each of the multiple different voice feedback recordings corresponds to a different style of music. Claim 2: The method of claim 1, wherein each of the multiple different voice feedback recordings of the different voice artists corresponds to a different style of music Claim 20: The system of claim 18, wherein the identifying musical characteristic comprises a particular style of music selected from a predefined list of different styles of music. Claim 2: and wherein the identifying musical characteristic comprises a particular style of music selected from a predefined list of different styles of music. Claim 21: The system of claim 18, wherein the memory further stores instructions for, before the storing step: receiving a first voice recording; and generating a first set of multiple voice recordings from the first voice recording using at least one of machine learning or natural language generation. Claim 3: The method of claim 2, further comprising, before the storing step: receiving a first voice recording from a first voice artist; and generating a first set of multiple voice recordings from the first voice recording, using at least one of machine learning or natural language generation. Note# 1 Although the claims at issue are not identical, they are not patentably distinct from each other because simply changing the statutory category from computer program product and system, to a method and removing inherent and/or unnecessary limitations/step would be within the level of one of ordinary skill in the art. It is well settled that the omission of an element, e.g. “a computer-readable storage medium”, and its function is an obvious expedient if the remaining elements perform the same function as before. In re Karlson, 136 USPQ 184 (CCPA 1963). Also note Ex parte Rainu, 168 USPQ 375 (Bd. App. 1969). Omission of a reference element or step whose function is not needed would be obvious to one of ordinary skill in the art. Claims 2-21 are rejected on the ground of nonstatutory double patenting as being unpatentable over claim 1 of U.S. Patent No. US 12283271 B2. Although the claims at issue are not identical, they are not patentably distinct from each other. See the table below for mapping of similar limitations. Instant Application U.S. Patent No. US 12283271 B2 Claim 2: A method of providing voice feedback, comprising: Claim 1: A media delivery system comprising: a processing device; and a memory device coupled to the processing device and storing instructions that, when executed by the processing device, cause the media delivery system to: See Note #1 above storing multiple different voice feedback recordings in at least one computer-readable storage device; store the plurality of sets of voice feedback recordings See Note #1 above receiving a listener command corresponding to a musical selection; receive, from a media playback device located in a vehicle, a command received as input from a user of the media playback device, the command associated with playback of a media content item; determining, with a processing device, an identifying musical characteristic of the musical selection; in response to receiving the command: determine an action to be performed in response to the command; pair a set of voice feedback recordings, from the plurality of sets of voice feedback recordings generated from the at least one initial voice feedback recording using the machine learning model, with the media content item based on a characteristic of the media content item; selecting a first voice feedback recording from the multiple different voice feedback recordings, using the processing device, wherein the first voice feedback recording corresponds to the identifying musical characteristic; with the media content item based on a characteristic of the media content item; select, from the set of voice feedback recordings paired with the media content item, a voice feedback recording that corresponds to the determined action; and causing playback of the first voice feedback recording and the musical selection via a media playback system. and provide an instruction causing playback of the voice feedback recording to the media playback device in the vehicle. Claims 3-16 See Note #1 above Claim 17: A non-transitory computer readable medium for use on a computer system containing computer-executable programming instructions, the instructions including instructions for: Claim 1: A media delivery system comprising: a processing device; and a memory device coupled to the processing device and storing instructions that, when executed by the processing device, cause the media delivery system to: See Note #1 above storing multiple different voice feedback recordings in at least one computer-readable storage device; store the plurality of sets of voice feedback recordings See Note #1 above receiving a listener command corresponding to a musical selection; receive, from a media playback device located in a vehicle, a command received as input from a user of the media playback device, the command associated with playback of a media content item; determining, with a processing device, an identifying musical characteristic of the musical selection; in response to receiving the command: determine an action to be performed in response to the command; pair a set of voice feedback recordings, from the plurality of sets of voice feedback recordings generated from the at least one initial voice feedback recording using the machine learning model, with the media content item based on a characteristic of the media content item; selecting a first voice feedback recording from the multiple different voice feedback recordings, using the processing device, wherein the first voice feedback recording corresponds to the identifying musical characteristic; with the media content item based on a characteristic of the media content item; select, from the set of voice feedback recordings paired with the media content item, a voice feedback recording that corresponds to the determined action; and causing playback of the first voice feedback recording and the musical selection via a media playback system. and provide an instruction causing playback of the voice feedback recording to the media playback device in the vehicle. Claim 18: A system, comprising: one or more processors: memory storing instructions for execution by the one or more processors, including instructions for: Claim 1: A media delivery system comprising: a processing device; and a memory device coupled to the processing device and storing instructions that, when executed by the processing device, cause the media delivery system to: See Note #1 above storing multiple different voice feedback recordings in at least one computer-readable storage device; store the plurality of sets of voice feedback recordings See Note #1 above receiving a listener command corresponding to a musical selection; receive, from a media playback device located in a vehicle, a command received as input from a user of the media playback device, the command associated with playback of a media content item; determining, with a processing device, an identifying musical characteristic of the musical selection; in response to receiving the command: determine an action to be performed in response to the command; pair a set of voice feedback recordings, from the plurality of sets of voice feedback recordings generated from the at least one initial voice feedback recording using the machine learning model, with the media content item based on a characteristic of the media content item; selecting a first voice feedback recording from the multiple different voice feedback recordings, using the processing device, wherein the first voice feedback recording corresponds to the identifying musical characteristic; with the media content item based on a characteristic of the media content item; select, from the set of voice feedback recordings paired with the media content item, a voice feedback recording that corresponds to the determined action; and causing playback of the first voice feedback recording and the musical selection via a media playback system. and provide an instruction causing playback of the voice feedback recording to the media playback device in the vehicle. Claims 19-21 See Note #1 above Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 2-4, 8-9 and 17-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Lindahl et al. (US 20110066438 A1). Regarding Claim 2: Lindahl teaches a method of providing voice feedback, comprising: storing multiple different voice feedback recordings in at least one computer-readable storage device(Para [0063], Ln 1-12, voice feedback items may be previously recorded by a recording artist and associated with a primary media item…enhanced media file is played back on either the host device. Para [0066], Ln 1-6, media files may include a user's favorite songs, an entire album by a recording artist); receiving a listener command corresponding to a musical selection(Para [0073], Ln 13-21, feedback event may be a track change or playlist change that is manually initiated by a user or automatically initiated by a media player application……feedback event may occur on demand by a user. For instance, the media player application may provide a command that the user may select….in order to hear voice feedback.); determining, with a processing device, an identifying musical characteristic of the musical selection(Para [0063], Ln 1-12, voice feedback items may be previously recorded by a recording artist and associated with a primary media item…enhanced media file is played back on either the host device…..the pre-associated voice feedback data may be concurrently played, thereby allowing a user to listen to a voice feedback announcement (e.g., artist, track, album, etc.) or commentary that is spoken by the recording artist. Para [0073], Ln 13-21, feedback event may be a track change or playlist change that is manually initiated by a user or automatically initiated by a media player application); selecting a first voice feedback recording from the multiple different voice feedback recordings, using the processing device, wherein the first voice feedback recording corresponds to the identifying musical characteristic(Para [0063], Ln 1-12, voice feedback items may be previously recorded by a recording artist and associated with a primary media item…enhanced media file is played back on either the host device…..the pre-associated voice feedback data may be concurrently played, thereby allowing a user to listen to a voice feedback announcement (e.g., artist, track, album, etc.) or commentary that is spoken by the recording artist. Para [0073], Ln 13-21, feedback event may be a track change or playlist change that is manually initiated by a user or automatically initiated by a media player application); and causing playback of the first voice feedback recording and the musical selection via a media playback system(Para [0059], Ln 7-9, upon playback of the enhanced media item, secondary media data may be played concurrently). Regarding Claim 3: Lindahl teaches the method of claim 2, wherein each of the multiple different voice feedback recordings corresponds to a different style of music(Para [0086], Ln 1-13, Based on the determined genre, the method 180 may then include varying characteristics of a voiceover announcement (or other audio feedback) based on the identified genre). Regarding Claim 4: Lindahl teaches the method of claim 2, wherein the identifying musical characteristic comprises a particular style of music selected from a predefined list of different styles of music(See decision tree in Fig 10. Para [0086], Ln 1-13, Based on the determined genre, the method 180 may then include varying characteristics of a voiceover announcement (or other audio feedback) based on the identified genre). Regarding Claim 8: Lindahl teaches the method of claim 2, wherein storing the multiple different voice feedback recordings comprises storing different tempo recordings(Para [0009], Ln 1-11, voice feedback data may then be processed to vary one or more characteristics of the voice feedback data based on the one or more parameters determined from the audio data. Voice feedback characteristics that may be varied through such processing may include pitch, tempo, reverberation, mono or stereo imaging, timbre, equalization, and volume, among others. Particularly, in some embodiments, the variation of voice feedback characteristics may provide facilitate better integration of the voice feedback with the primary audio data with which it is associated. Para [0070], Ln 10-20, For instance, in some embodiments, the parameters on which the alteration of the voiceover announcement is based may include one or more of a reverberation parameter, a timbre parameter, a pitch parameter, a volume parameter, an equalization parameter, a tempo parameter). Regarding Claim 9: Lindahl teaches the method of claim 2, wherein receiving the listener command comprises receiving at least one of a shelf selection or a card selection(Para [0033], Ln 6-21, For example, one or more of the user input structures may include a wheel structure that may allow a user to select various icons 30 displayed by the GUI 28. Additionally, the icons 30 may also be selected via the touch screen interface of the display 24.(shelf). Para [0065], Ln 1-9, user may conveniently select a defined playlist to load the entire group of media files without having to specify the location of each media file.(card). Para [0067], Ln 1-13, user assigned the name "Favorite Songs" to the defined playlist, a voice synthesis program may create and associate a secondary media item with playlist, such that when the playlist is loaded by the media player application or when a media item from the playlist is initially played, the secondary media item may be played back concurrently and announce the name of the playlist as "Favorite Songs."). Regarding Claim 17: Lindahl teaches a non-transitory computer readable medium for use on a computer system containing computer-executable programming instructions, the instructions including instructions for(Para [0041], Ln 1-7, may include a volatile memory, such as RAM. Para [0040], Ln 1-8, The operation of the device 10 may be generally controlled by one or more processors 50, which may provide the processing capability required to execute an operating system): storing multiple different voice feedback recordings in at least one computer-readable storage device(Para [0063], Ln 1-12, voice feedback items may be previously recorded by a recording artist and associated with a primary media item…enhanced media file is played back on either the host device. Para [0066], Ln 1-6, media files may include a user's favorite songs, an entire album by a recording artist); receiving a listener command corresponding to a musical selection(Para [0073], Ln 13-21, feedback event may be a track change or playlist change that is manually initiated by a user or automatically initiated by a media player application……feedback event may occur on demand by a user. For instance, the media player application may provide a command that the user may select….in order to hear voice feedback.); determining, with a processing device, an identifying musical characteristic of the musical selection(Para [0063], Ln 1-12, voice feedback items may be previously recorded by a recording artist and associated with a primary media item…enhanced media file is played back on either the host device…..the pre-associated voice feedback data may be concurrently played, thereby allowing a user to listen to a voice feedback announcement (e.g., artist, track, album, etc.) or commentary that is spoken by the recording artist. Para [0073], Ln 13-21, feedback event may be a track change or playlist change that is manually initiated by a user or automatically initiated by a media player application); selecting a first voice feedback recording from the multiple different voice feedback recordings, using the processing device, wherein the first voice feedback recording corresponds to the identifying musical characteristic(Para [0063], Ln 1-12, voice feedback items may be previously recorded by a recording artist and associated with a primary media item…enhanced media file is played back on either the host device…..the pre-associated voice feedback data may be concurrently played, thereby allowing a user to listen to a voice feedback announcement (e.g., artist, track, album, etc.) or commentary that is spoken by the recording artist. Para [0073], Ln 13-21, feedback event may be a track change or playlist change that is manually initiated by a user or automatically initiated by a media player application); and causing playback of the first voice feedback recording and the musical selection via a media playback system(Para [0059], Ln 7-9, upon playback of the enhanced media item, secondary media data may be played concurrently). Regarding Claim 18: Lindahl teaches a system, comprising: one or more processors: memory storing instructions for execution by the one or more processors, including instructions for(Para [0041], Ln 1-7, may include a volatile memory, such as RAM. Para [0040], Ln 1-8, The operation of the device 10 may be generally controlled by one or more processors 50, which may provide the processing capability required to execute an operating system): storing multiple different voice feedback recordings in at least one computer-readable storage device(Para [0063], Ln 1-12, voice feedback items may be previously recorded by a recording artist and associated with a primary media item…enhanced media file is played back on either the host device. Para [0066], Ln 1-6, media files may include a user's favorite songs, an entire album by a recording artist); receiving a listener command corresponding to a musical selection(Para [0073], Ln 13-21, feedback event may be a track change or playlist change that is manually initiated by a user or automatically initiated by a media player application……feedback event may occur on demand by a user. For instance, the media player application may provide a command that the user may select….in order to hear voice feedback.); determining, with a processing device, an identifying musical characteristic of the musical selection(Para [0063], Ln 1-12, voice feedback items may be previously recorded by a recording artist and associated with a primary media item…enhanced media file is played back on either the host device…..the pre-associated voice feedback data may be concurrently played, thereby allowing a user to listen to a voice feedback announcement (e.g., artist, track, album, etc.) or commentary that is spoken by the recording artist. Para [0073], Ln 13-21, feedback event may be a track change or playlist change that is manually initiated by a user or automatically initiated by a media player application); selecting a first voice feedback recording from the multiple different voice feedback recordings, using the processing device, wherein the first voice feedback recording corresponds to the identifying musical characteristic(Para [0063], Ln 1-12, voice feedback items may be previously recorded by a recording artist and associated with a primary media item…enhanced media file is played back on either the host device…..the pre-associated voice feedback data may be concurrently played, thereby allowing a user to listen to a voice feedback announcement (e.g., artist, track, album, etc.) or commentary that is spoken by the recording artist. Para [0073], Ln 13-21, feedback event may be a track change or playlist change that is manually initiated by a user or automatically initiated by a media player application); and causing playback of the first voice feedback recording and the musical selection via a media playback system(Para [0059], Ln 7-9, upon playback of the enhanced media item, secondary media data may be played concurrently). Regarding Claim 19: Claim 19 contains similar limitations as Claim 3, as is therefore rejected for the same reasons. Regarding Claim 20: Claim 20 contains similar limitations as Claim 4, as is therefore rejected for the same reasons. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 5-6 and 21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lindahl as applied to claim 2 above, and further in view of Lowe (US 7142645 B2). Regarding Claim 5: Lindahl teaches the method of claim 2, but does not teach further comprising, before the storing step: receiving a first voice recording; and generating a first set of multiple voice recordings from the first voice recording using at least one of machine learning or natural language generation. In the same field of Speech interfaces, Lowe teaches further comprising, before the storing step: receiving a first voice recording(Col 12, Ln 39-52, Alternatively, the system may utilize a plurality of master clips that each begin and/or end at the point where an insert clip is to be placed. When the master clips are merged together with one or more appropriate insert clips the result is a seamless media clip ready for playback); and generating a first set of multiple voice recordings from the first voice recording using at least one of machine learning or natural language generation(Col 12, Ln 39-52, Alternatively, the system may utilize a plurality of master clips that each begin and/or end at the point where an insert clip is to be placed. When the master clips are merged together with one or more appropriate insert clips the result is a seamless media clip ready for playback. Col 13, Ln 4-24, The applications generated with embodiments of the invention reflect the flow of natural language. This is accomplished when a creator of the application writes at least one "generic" filler for every slot in the application and/or provides an alphabetic set of "generic" fillers for slots with highly variable information (e.g. name) and accounts for phonemic blending that occurs across closely annunciated phrases). It would have been obvious for one skilled in the art, at the effective time of filling, to modify Lindahl with the customized audio clips of Lowe, as it improves user experience by providing personalized audio clips(Col 1, Ln 45-60). Regarding Claim 6: The combination of Lindahl and Lowe teaches the method of claim 5, but does not teach wherein the first set of multiple voice recordings comprises at least one of different tempos, different words, different pitches or different speaking styles. In the same field of Speech interfaces, Lowe teaches wherein the first set of multiple voice recordings comprises at least one of different tempos, different words, different pitches or different speaking styles(Col 13, Ln 4-24, The applications generated with embodiments of the invention reflect the flow of natural language. This is accomplished when a creator of the application writes at least one "generic" filler for every slot in the application and/or provides an alphabetic set of "generic" fillers for slots with highly variable information (e.g. name) and accounts for phonemic blending that occurs across closely annunciated phrases). It would have been obvious for one skilled in the art, at the effective time of filling, to modify the combination of Lindahl and Lowe with the customized audio clips of Lowe, as it improves user experience by providing personalized audio clips(Col 1, Ln 45-60). Regarding Claim 21: Claim 21 contains similar limitations as Claim 5, as is therefore rejected for the same reasons. Claim(s) 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lindahl as applied to claim 2 above, and further in view of Mackenzie et al. (US 20130007617 A1). Regarding Claim 7: Lindahl teaches the method of claim 2, but does not teach wherein the musical selection is selected from the group consisting of a piece of music, an album, an artist, a style of music, a playlist, a shelf of music and a card of music(Specifically, Lindahl doesn’t teach all elements, and claim language “the group consisting of” is being interpreted as requiring all elements. see MPEP 2111.03, II consisting of, Para 2, A claim element defined by selection from a group of alternatives…. In the absence of such qualifying language there is a presumption that the Markush group is closed to combinations or mixtures). In the same field of media playback, Mackenzie teaches wherein the musical selection is selected from the group consisting of a piece of music(Para [0024], Ln 1-24, All songs), an album(Para [0024], Ln 1-24, albums), an artist(Para [0024], Ln 1-24, artists), a style of music(Para [0024], Ln 1-24, genres), a playlist(Para [0024], Ln 1-24, playlist), a shelf of music(Para [0024], Ln 1-24, All songs. Also list of songs in Fig 1) and a card of music(Para [0024], Ln 1-24. See Fig 1, Song 3, see set of information regarding song. Shelf and card are being interpreted under broadest reasonable interpretation). It would have been obvious for one skilled in the art, at the effective time of filing to modify Lindahl with the media playback system of Mackenzie, as it improves user convenience(Para [0003], Ln 1-18). Claim(s) 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lindahl as applied to claim 2 above, and further in view of Ales (US 20160239876 A1). Regarding Claim 10: Lindahl teaches the method of claim 2, but does not teach further comprising: creating a voice beat grid for the first voice feedback recording; and creating a music beat grid for the musical selection. In the same field of media playback, Ales teaches creating a voice beat grid for the first voice feedback recording; and creating a music beat grid for the musical selection(Para [0045], Ln 14-28, Music Information retrieval (MIR) processes may include bar/beat grid detection routines. Para [0010], Ln 1-20, decision engine for mixing a first song recording with a first non-song musical content item…..accessing metadata for one or more song recordings including a first song recording, and metadata for one or more non-song musical content items including a first non-song musical content item; responsive to playback of at least the first song recording, interpreting the metadata for the first song recording; identifying the first non-song musical content item for playback at or near an end of the first song recording based on a comparison of the metadata of the first song recording and the metadata of the first non-song musical content item; and in response to determining that the first song recording is at or near its end of playback, forming an altered playback of the first non-song musical content item by creating an alteration of the first non-song musical content item to be rhythmically continuous in terms of tempo to the first song recording. Para [0177], Ln 1-25, non-song musical content may be inserted as part of a segue. This non-song musical content may consist of advertising with backing music or of short musical/sound logos known as “mnemonics” (e.g., the “Intel Inside” musical figure). Such musical/sound logos may serve to brand the music service licensee (e.g., a station ID “button”). For example, the playback module 310 may include instructions that at a particular time, or after a particular number of songs have played, the segue between songs would be to such non-song musical content. Such content could overlap the end portion of the current song (as is the case with song-to-song segues) or begin immediately and contiguously at the end of the current song in such a way as to be rhythmically continuous in terms of tempo (as related to the current song). Para [0062], Ln 1-10, runtime segue selection performs the steps of evaluating a song pair based upon whether the song pair is tempo-discrete (e.g., failed the temporal evaluation), tempo-concurrent). It would have been obvious for one skilled in the art, at the effective time of filling, to modify Lindahl, with the tempo synchronization of Ales, as it improves the user experience by making the playback of additional content more cohesive(Para [0009], Ln 10-23). Claim(s) 11-12 and 14-15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lindahl as applied to claim 2 above, and further in view of Otto et al. (US 20150106394 A1). Regarding Claim 11: Lindahl teaches the method of claim 2, but does not teach wherein the first voice feedback recording is played at least partially before the musical selection is played by the media playback system. In the same field of media playback, Otto teaches wherein the first voice feedback recording is played at least partially before the musical selection is played by the media playback system(Para [0037], Ln 1-14, song announcements are made can be varied to make the system more playful. In one example, the song information might be announced before and after each song. For example, the audio snippet might say the following: "That was "Billie Jean" by Michael Jackson. Up next is "Satisfaction" by the Rolling Stones." In another example, instead of playing an audio snippet during a pause between songs, the audio insertion logic 112 could fade in the audio snippet, e.g., by causing the volume of the song being played to be lowered and playing the audio snippet over the song being played back. After the audio snippet has been played, the audio insertion logic could cause the volume to be increased back to the original level). It would have been obvious for one skilled in the art, at the effective time of filling, to modify Lindahl with the audio announcement system of Otto, as it improves user convenience and safety(Para [0001], Ln 1-9). Regarding Claim 12: The combination of Lindahl and Otto teaches the method of claim 11, but does not teach wherein a portion of the first voice feedback recording is played at a same time that a beginning portion of the musical selection is played. In the same field of media playback, Otto teaches wherein a portion of the first voice feedback recording is played at a same time that a beginning portion of the musical selection is played(Para [0037], Ln 1-14, song announcements are made can be varied to make the system more playful. In one example, the song information might be announced before and after each song. For example, the audio snippet might say the following: "That was "Billie Jean" by Michael Jackson. Up next is "Satisfaction" by the Rolling Stones." In another example, instead of playing an audio snippet during a pause between songs, the audio insertion logic 112 could fade in the audio snippet, e.g., by causing the volume of the song being played to be lowered and playing the audio snippet over the song being played back. After the audio snippet has been played, the audio insertion logic could cause the volume to be increased back to the original level). It would have been obvious for one skilled in the art, at the effective time of filling, to modify the combination of Lindahl and Otto with the audio announcement system of Otto, as it improves user convenience and safety(Para [0001], Ln 1-9). Regarding Claim 14: Lindahl teaches the method of claim 2, but does not teach wherein the first voice feedback recording is played at least partially after the musical selection is played by the media playback system. In the same field of media playback, Otto teaches wherein the first voice feedback recording is played at least partially after the musical selection is played by the media playback system(Para [0037], Ln 1-14, song announcements are made can be varied to make the system more playful. In one example, the song information might be announced before and after each song. For example, the audio snippet might say the following: "That was "Billie Jean" by Michael Jackson. Up next is "Satisfaction" by the Rolling Stones." In another example, instead of playing an audio snippet during a pause between songs, the audio insertion logic 112 could fade in the audio snippet, e.g., by causing the volume of the song being played to be lowered and playing the audio snippet over the song being played back. After the audio snippet has been played, the audio insertion logic could cause the volume to be increased back to the original level). It would have been obvious for one skilled in the art, at the effective time of filling, to modify Lindahl with the audio announcement system of Otto, as it improves user convenience and safety(Para [0001], Ln 1-9). Regarding Claim 15: Lindahl teaches the method of claim 2, but does not teach wherein the multiple different voice feedback recordings comprise multiple introductions of multiple possible musical selections. In the same field of media playback, Otto teaches wherein the multiple different voice feedback recordings comprise multiple introductions of multiple possible musical selections(Para [0030], Ln 1-19, queue 120 being streamed by steaming service 116 includes song A (intro), song A (music), song B (intro), song B (music), song C (intro), and song C (music). In this embodiment, the "intro" for each of songs A-C is the audio snippet generated by the text-to-voice service 110 using the metadata 104a for each song. The "music" for each of songs A-C is the audio data 104b contained in each song file 104. The queue 120' is the queue of songs that is displayed on the screen of client device 118. As shown in FIG. 2A, the queue 120' includes only songs A-C. In other words, the "intro" for each of songs A-C is not made visible to the user on the client device). It would have been obvious for one skilled in the art, at the effective time of filling, to modify Lindahl with the audio announcement system of Otto, as it improves user convenience and safety(Para [0001], Ln 1-9). Claim(s) 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Lindahl and Otto as applied to claim 12 above, and further in view of Ales. Regarding Claim 13: The combination of Lindahl and Otto teaches the method of claim 12, but does not teach wherein at least the portion of the first voice feedback recording is played on-beat with the musical selection. In the same field of media playback, Ales teaches wherein at least the portion of the first voice feedback recording is played on-beat with the musical selection(Para [0045], Ln 14-28, Music Information retrieval (MIR) processes may include bar/beat grid detection routines. Para [0010], Ln 1-20, decision engine for mixing a first song recording with a first non-song musical content item…..accessing metadata for one or more song recordings including a first song recording, and metadata for one or more non-song musical content items including a first non-song musical content item; responsive to playback of at least the first song recording, interpreting the metadata for the first song recording; identifying the first non-song musical content item for playback at or near an end of the first song recording based on a comparison of the metadata of the first song recording and the metadata of the first non-song musical content item; and in response to determining that the first song recording is at or near its end of playback, forming an altered playback of the first non-song musical content item by creating an alteration of the first non-song musical content item to be rhythmically continuous in terms of tempo to the first song recording. Para [0177], Ln 1-25, non-song musical content may be inserted as part of a segue. This non-song musical content may consist of advertising with backing music or of short musical/sound logos known as “mnemonics” (e.g., the “Intel Inside” musical figure). Such musical/sound logos may serve to brand the music service licensee (e.g., a station ID “button”). For example, the playback module 310 may include instructions that at a particular time, or after a particular number of songs have played, the segue between songs would be to such non-song musical content. Such content could overlap the end portion of the current song (as is the case with song-to-song segues) or begin immediately and contiguously at the end of the current song in such a way as to be rhythmically continuous in terms of tempo (as related to the current song). Para [0062], Ln 1-10, runtime segue selection performs the steps of evaluating a song pair based upon whether the song pair is tempo-discrete (e.g., failed the temporal evaluation), tempo-concurrent). It would have been obvious for one skilled in the art, at the effective time of filling, to modify the combination of Lindahl and Otto, with the tempo synchronization of Ales, as it improves the user experience by making the playback of additional content more cohesive(Para [0009], Ln 10-23). Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lindahl as applied to claim 2 above, and further in view of Rohra (WO 2011108007 A2). Regarding Claim 16: Lindahl teaches the method of claim 2, but does not teach further comprising customizing at least the first voice feedback recording to address a listener by name. In the same field of media playback, Rohra teaches further comprising customizing at least the first voice feedback recording to address a listener by name(Pg 9, Para 5, Ln 1-8, messages will include names of the subscribers. Pg 1, abstract, Ln 1-5, The system enables creation of ready to use personalized multimedia content having name, personalized greeting and high quality animation, music and graphics. The content can be in the form of multimedia messages). It would have been obvious for one skilled in the art, at the effective time of filling, to modify Lindahl with the personalized content system of Rohra, as the personalized content can improve user experience(Pg 3, Para 1, Ln 1-6. Pg 3, Para 3, Ln 1-5). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See IDS filed on 04/28/2025 for references that are pertinent to the disclosure, made of record during examination of parent applications. Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDER G MARLOW whose telephone number is (571)272-4536. The examiner can normally be reached Monday - Thursday 10:00 am - 8:00 pm EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richmond Dorvil can be reached at (571)272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /ALEXANDER G MARLOW/ Assistant Examiner, Art Unit 2658 /RICHEMOND DORVIL/ Supervisory Patent Examiner, Art Unit 2658
Read full office action

Prosecution Timeline

Jan 16, 2025
Application Filed
Jul 28, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694204
SYSTEM AND METHOD FOR COMPARING DOCUMENTS USING AN ARTIFICIAL INTELLIGENCE (AI) MODEL
2y 4m to grant Granted Jul 28, 2026
Patent 12682903
Voice Query QoS based on Client-Computed Content Metadata
2y 9m to grant Granted Jul 14, 2026
Patent 12670896
INFORMATION PROCESSING METHOD, NON-TRANSITORY RECORDING MEDIUM, INFORMATION PROCESSING APPARATUS, AND INFORMATION PROCESSING SYSTEM
3y 6m to grant Granted Jun 30, 2026
Patent 12664992
SPATIAL AUDIO PARAMETER ENCODING AND ASSOCIATED DECODING
3y 3m to grant Granted Jun 23, 2026
Patent 12646515
SELECTIVELY PROVIDING ENHANCED CLARIFICATION PROMPTS IN AUTOMATED ASSISTANT INTERACTIONS
2y 8m to grant Granted Jun 02, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
79%
Grant Probability
97%
With Interview (+18.1%)
2y 8m (~1y 1m remaining)
Median Time to Grant
Low
PTA Risk
Based on 84 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month