Prosecution Insights
Last updated: August 17, 2026
Application No. 18/970,443

AUDIO PROCESSING DEVICE, AUDIO PROCESSING METHOD, AND COMPUTER PROGRAM PRODUCT

Non-Final OA §103
Filed
Dec 05, 2024
Priority
Feb 21, 2024 — JP 2024-024305
Examiner
BECKER, TYLER JUSTIN
Art Unit
Tech Center
Assignee
Panasonic Holdings Corporation
OA Round
1 (Non-Final)
74%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
80%
With Interview

Examiner Intelligence

Grants 74% — above average
74%
Career Allowance Rate
17 granted / 23 resolved
+13.9% vs TC avg
Moderate +6% lift
Without
With
+6.3%
Interview Lift
resolved cases with interview
Typical timeline
2y 8m
Avg Prosecution
15 currently pending
Career history
45
Total Applications
across all art units

Statute-Specific Performance

§101
19.2%
-20.8% vs TC avg
§103
51.1%
+11.1% vs TC avg
§102
14.3%
-25.7% vs TC avg
§112
15.4%
-24.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 23 resolved cases

Office Action

§103
DETAILED ACTION This action is in response to the application filed on December 5th, 2024. Claims 1-19 are pending and have been examined. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1, 2, 5, 6, 9, 10, 13, 14, and 17-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Shinto et al. (JP Pat. No. 2009086132 A hereinafter Shinto), in view of Park et al. (KR Pat. No. 20200040425 A hereinafter Park). Regarding claim 1, Shinto discloses an audio processing device comprising: a memory in which a program is stored; and a processor coupled to the memory and configured to perform processing by executing the program, the processing including (Shinto, Pg. 7, paragraph 11: " The registration unit 101, the reception unit 102, the voice recognition unit 103, the control unit 104, the output unit 105, the setting unit 106, the change unit 107, and the input unit 108 included in the voice recognition device 100 illustrated in FIG. Means that the CPU 301 executes a predetermined program using the programs and data recorded in the ROM 302, RAM 303, magnetic disk 305, optical disk 307, etc. in the navigation device 300 shown in FIG. To realize its function."): calculating utterance history coefficients for the users based on utterance histories of the users (Shinto, Pg. 10, paragraph 6: "According to the above-described processing, the voice information of the user with high priority is recognized among the received voices, and the voice information other than the user with low priority is removed. It is possible to prevent misrecognition by voice other than the person's utterance. In particular, in the navigation device 300, as a user who speaks frequently, a user with high driving frequency is targeted, and various setting information such as route search conditions and search conditions associated with the user is read. Therefore, it is possible to save the user from having to select various setting information corresponding to the user."); selecting a predetermined number of users from among the users as recognition targets of the pieces of the audio information based on the utterance history coefficients (Shinto, Pg. 4, paragraph 7: "For example, when 10 users are registered, the priority indicates a value of 10 levels. This configuration recognizes the voice of the user with the higher priority. For example, when the voices of the users with the fifth and eighth priorities are received, the voice of the user with the fifth highest priority is received. The target of speech recognition. In addition, when a user with the highest priority is set as a person to be recognized and the voice of the user with the highest priority is received, the voice of the user with the highest priority is recognized and prioritized. The setting of the first-ranked user may be changed as a recognition target person."); outputting recognition target person information of the selected recognition targets (Shinto, Pg. 4, paragraph 2: "The setting unit 106 is set with a user (hereinafter referred to as a “recognition target person”) as a voice recognition target from the registration unit 101 in which voice information of a plurality of users is registered."; Here, the recognition target is seen as being output from the registration unit to the setting unit.). However, Shinto fails to expressly recite calculating, from registration information in which pieces of audio information of a plurality of users are registered, group information for each user group in which features of the pieces of the audio information are similar between the users; and outputting the recognition target person information and information indicating that a plurality of users having a same user group are included when the plurality of users having the same user group are included in the selected recognition targets. Park teaches calculating, from registration information in which pieces of audio information of a plurality of users are registered, group information for each user group in which features of the pieces of the audio information are similar between the users (Park, Pg. 6, paragraph 2: "The reference score setting unit 350 may determine whether the corresponding users are a general group or a special group based on the similarity of voices between pre-registered home users. That is, when the voice similarity between the users in the house exceeds a reference score (hereinafter referred to as 'first reference score' for convenience of description) for determining whether a special group is present, the reference score setting unit 350 determines the corresponding users It can be classified (recognized) as a 'special group'. On the other hand, if the voice similarity between the users in the home is less than or equal to the first reference score, the reference score setting unit 350 may classify (recognize) the users as a 'normal group'."); and outputting the recognition target person information and information indicating that a plurality of users having a same user group are included when the plurality of users having the same user group are included in the selected recognition targets (Park, Pg. 7, paragraph 6: "On the other hand, as a result of the check in step 460, when the measured similarity score is greater than the first reference score, the speaker recognition device 300 recognizes a user requesting voice registration and a pre-registered home user as a special group and is a speaker The reference score for identification may be updated (S480). At this time, the speaker recognition device 300 may update the reference score for speaker identification using the measured similarity score."; Here, the indication of a special group is seen as an output that is then used to update the reference score.). Shinto and Park are analogous arts because they each belong to the same field of speech recognition systems. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition device of Shinto to incorporate the teachings of Park to create user groups based on audio similarities. This ensures that the system can accurately identify a speaker even if other speakers have similar voices (Park, Pg. 2, paragraph 6). As such, the system works better overall, and the user experience is improved. Regarding claim 2, the rejection of claim 1 is incorporated. Shinto, in view of Park, discloses all of the elements of the claimed invention as stated above. Shinto further discloses wherein the processing includes calculating a value of the utterance history coefficient to be higher for a user corresponding to a piece of the audio information having a larger number of times of recognition as an utterance in past, among the registered pieces of the audio information of the users (Shinto, Pg. 10, paragraph 6: "In particular, in the navigation device 300, as a user who speaks frequently, a user with high driving frequency is targeted, and various setting information such as route search conditions and search conditions associated with the user is read."). Regarding claim 5, the rejection of claim 1 is incorporated. Shinto, in view of Park, discloses all of the elements of the claimed invention as stated above. Shinto further discloses wherein the processing includes deleting, from the recognition targets, a user whose utterance history coefficient satisfies a predetermined condition among the plurality of users having the same user group, based on the utterance history coefficients of the plurality of users having the same user group (Shinto, Pg. 10, paragraph 6: " According to the above-described processing, the voice information of the user with high priority is recognized among the received voices, and the voice information other than the user with low priority is removed."). However, Shinto fails to expressly recite when the plurality of users having the same user group are included in the recognition targets to be selected. Park further teaches when the plurality of users having the same user group are included in the recognition targets to be selected (Park, Pg. 6, paragraph 2: "The reference score setting unit 350 may determine whether the corresponding users are a general group or a special group based on the similarity of voices between pre-registered home users. That is, when the voice similarity between the users in the house exceeds a reference score (hereinafter referred to as 'first reference score' for convenience of description) for determining whether a special group is present, the reference score setting unit 350 determines the corresponding users It can be classified (recognized) as a 'special group'. On the other hand, if the voice similarity between the users in the home is less than or equal to the first reference score, the reference score setting unit 350 may classify (recognize) the users as a 'normal group'."). The same motivation for claim 1 applies equally to claim 5. Regarding claim 6, the rejection of claim 2 is incorporated. Shinto, in view of Park, discloses all of the elements of the claimed invention as stated above. Shinto further discloses wherein the processing includes deleting, from the recognition targets, a user whose utterance history coefficient satisfies a predetermined condition among the plurality of users having the same user group, based on the utterance history coefficients of the plurality of users having the same user group (Shinto, Pg. 10, paragraph 6: " According to the above-described processing, the voice information of the user with high priority is recognized among the received voices, and the voice information other than the user with low priority is removed."). However, Shinto fails to expressly recite when the plurality of users having the same user group are included in the recognition targets to be selected. Park further teaches when the plurality of users having the same user group are included in the recognition targets to be selected (Park, Pg. 6, paragraph 2: "The reference score setting unit 350 may determine whether the corresponding users are a general group or a special group based on the similarity of voices between pre-registered home users. That is, when the voice similarity between the users in the house exceeds a reference score (hereinafter referred to as 'first reference score' for convenience of description) for determining whether a special group is present, the reference score setting unit 350 determines the corresponding users It can be classified (recognized) as a 'special group'. On the other hand, if the voice similarity between the users in the home is less than or equal to the first reference score, the reference score setting unit 350 may classify (recognize) the users as a 'normal group'."). The same motivation for claim 1 applies equally to claim 6. Regarding claim 9, the rejection of claim 1 is incorporated. Shinto, in view of Park, discloses all of the elements of the claimed invention as stated above. Shinto further discloses wherein the processing further includes receiving, from a user of the users, selection of addition or deletion of a recognition target person to or from the recognition target person information (Shinto, Pg. 4, paragraph 4: "In this case, the registration unit 101 registers the user's voice information and a predetermined vocabulary for changing the recognition target person. The voice recognition unit 103 recognizes the voice information of the user registered in the registration unit 101 and a predetermined vocabulary among the voices received by the reception unit 102. Further, the changing unit 107 changes the recognition target person set in the setting unit 106 to the user who spoke, based on the result recognized by the voice recognition unit 103."). Regarding claim 10, the rejection of claim 2 is incorporated. Shinto, in view of Park, discloses all of the elements of the claimed invention as stated above. Shinto further discloses wherein the processing further includes receiving, from a user of the users, selection of addition or deletion of a recognition target person to or from the recognition target person information (Shinto, Pg. 4, paragraph 4: "In this case, the registration unit 101 registers the user's voice information and a predetermined vocabulary for changing the recognition target person. The voice recognition unit 103 recognizes the voice information of the user registered in the registration unit 101 and a predetermined vocabulary among the voices received by the reception unit 102. Further, the changing unit 107 changes the recognition target person set in the setting unit 106 to the user who spoke, based on the result recognized by the voice recognition unit 103."). Regarding claim 13, the rejection of claim 1 is incorporated. Shinto, in view of Park, discloses all of the elements of the claimed invention as stated above. Shinto further discloses wherein the processing further includes: inputting a voice of a speaker; and recognizing the speaker of the input voice based on the pieces of the audio information respectively corresponding to recognition target persons of the recognition targets (Shinto, Pg. 4, paragraph 2: " In this case, the voice recognition unit 103 recognizes the voice information of the recognition target person set in the setting unit 106 among the voices received by the reception unit 102. In this configuration, even when voice information of a plurality of users is registered in the registration unit 101, it is possible to recognize only the voice of the person to be recognized by setting."). Regarding claim 14 the rejection of claim 2 is incorporated. Shinto, in view of Park, discloses all of the elements of the claimed invention as stated above. Shinto further discloses wherein the processing further includes: inputting a voice of a speaker; and recognizing the speaker of the input voice based on the pieces of the audio information respectively corresponding to recognition target persons of the recognition targets (Shinto, Pg. 4, paragraph 2: " In this case, the voice recognition unit 103 recognizes the voice information of the recognition target person set in the setting unit 106 among the voices received by the reception unit 102. In this configuration, even when voice information of a plurality of users is registered in the registration unit 101, it is possible to recognize only the voice of the person to be recognized by setting."). Regarding claim 17, the rejection of claim 1 is incorporated. Shinto, in view of Park, discloses all of the elements of the claimed invention as stated above. Shinto further discloses wherein the processing further includes calculating the audio information based on an audio signal of the voice of the speaker and registering the audio information in the registration information (Shinto, Pg. 4, paragraph 6: " In the present embodiment, the registration unit 101 may register voice information of a plurality of users and information related to the priority associated with the voice information of the users and for identifying the recognition target person."). Regarding claim 18, Shinto discloses an audio processing method comprising: calculating utterance history coefficients for the users based on utterance histories of the users (Shinto, Pg. 10, paragraph 6: "According to the above-described processing, the voice information of the user with high priority is recognized among the received voices, and the voice information other than the user with low priority is removed. It is possible to prevent misrecognition by voice other than the person's utterance. In particular, in the navigation device 300, as a user who speaks frequently, a user with high driving frequency is targeted, and various setting information such as route search conditions and search conditions associated with the user is read. Therefore, it is possible to save the user from having to select various setting information corresponding to the user."); selecting a predetermined number of users from among the users as recognition targets of the pieces of the audio information based on the utterance history coefficients (Shinto, Pg. 4, paragraph 7: "For example, when 10 users are registered, the priority indicates a value of 10 levels. This configuration recognizes the voice of the user with the higher priority. For example, when the voices of the users with the fifth and eighth priorities are received, the voice of the user with the fifth highest priority is received. The target of speech recognition. In addition, when a user with the highest priority is set as a person to be recognized and the voice of the user with the highest priority is received, the voice of the user with the highest priority is recognized and prioritized. The setting of the first-ranked user may be changed as a recognition target person."); outputting recognition target person information of the selected recognition targets (Shinto, Pg. 4, paragraph 2: "The setting unit 106 is set with a user (hereinafter referred to as a “recognition target person”) as a voice recognition target from the registration unit 101 in which voice information of a plurality of users is registered."; Here, the recognition target is seen as being output from the registration unit to the setting unit.). However, Shinto fails to expressly recite calculating, from registration information in which pieces of audio information of a plurality of users are registered, group information for each user group in which features of the pieces of the audio information are similar between users; and outputting the recognition target person information and information indicating that a plurality of users having a same user group are included when the plurality of users having the same user group are included in the selected recognition targets. Park teaches calculating, from registration information in which pieces of audio information of a plurality of users are registered, group information for each user group in which features of the pieces of the audio information are similar between users (Park, Pg. 6, paragraph 2: "The reference score setting unit 350 may determine whether the corresponding users are a general group or a special group based on the similarity of voices between pre-registered home users. That is, when the voice similarity between the users in the house exceeds a reference score (hereinafter referred to as 'first reference score' for convenience of description) for determining whether a special group is present, the reference score setting unit 350 determines the corresponding users It can be classified (recognized) as a 'special group'. On the other hand, if the voice similarity between the users in the home is less than or equal to the first reference score, the reference score setting unit 350 may classify (recognize) the users as a 'normal group'."); and outputting the recognition target person information and information indicating that a plurality of users having a same user group are included when the plurality of users having the same user group are included in the selected recognition targets (Park, Pg. 7, paragraph 6: "On the other hand, as a result of the check in step 460, when the measured similarity score is greater than the first reference score, the speaker recognition device 300 recognizes a user requesting voice registration and a pre-registered home user as a special group and is a speaker The reference score for identification may be updated (S480). At this time, the speaker recognition device 300 may update the reference score for speaker identification using the measured similarity score."; Here, the indication of a special group is seen as an output that is then used to update the reference score.). Shinto and Park are analogous arts because they each belong to the same field of speech recognition systems. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition device of Shinto to incorporate the teachings of Park to create user groups based on audio similarities. This ensures that the system can accurately identify a speaker even if other speakers have similar voices (Park, Pg. 2, paragraph 6). As such, the system works better overall, and the user experience is improved. Regarding claim 19, Shinto discloses a non-transitory computer readable medium in and on which programmed instructions are embodied and stored, wherein the instructions, when executed by a computer (Shinto, Pg. 7, paragraph 11: " The registration unit 101, the reception unit 102, the voice recognition unit 103, the control unit 104, the output unit 105, the setting unit 106, the change unit 107, and the input unit 108 included in the voice recognition device 100 illustrated in FIG. Means that the CPU 301 executes a predetermined program using the programs and data recorded in the ROM 302, RAM 303, magnetic disk 305, optical disk 307, etc. in the navigation device 300 shown in FIG. To realize its function."), cause the computer to perform: calculating utterance history coefficients for the users based on utterance histories of the users (Shinto, Pg. 10, paragraph 6: "According to the above-described processing, the voice information of the user with high priority is recognized among the received voices, and the voice information other than the user with low priority is removed. It is possible to prevent misrecognition by voice other than the person's utterance. In particular, in the navigation device 300, as a user who speaks frequently, a user with high driving frequency is targeted, and various setting information such as route search conditions and search conditions associated with the user is read. Therefore, it is possible to save the user from having to select various setting information corresponding to the user."); selecting a predetermined number of users from among the users as recognition targets of the pieces of the audio information based on the utterance history coefficients (Shinto, Pg. 4, paragraph 7: "For example, when 10 users are registered, the priority indicates a value of 10 levels. This configuration recognizes the voice of the user with the higher priority. For example, when the voices of the users with the fifth and eighth priorities are received, the voice of the user with the fifth highest priority is received. The target of speech recognition. In addition, when a user with the highest priority is set as a person to be recognized and the voice of the user with the highest priority is received, the voice of the user with the highest priority is recognized and prioritized. The setting of the first-ranked user may be changed as a recognition target person."); outputting recognition target person information of the selected recognition targets (Shinto, Pg. 4, paragraph 2: "The setting unit 106 is set with a user (hereinafter referred to as a “recognition target person”) as a voice recognition target from the registration unit 101 in which voice information of a plurality of users is registered."; Here, the recognition target is seen as being output from the registration unit to the setting unit.). However, Shinto fails to expressly recite calculating, from registration information in which pieces of audio information of a plurality of users are registered, group information for each user group in which features of the pieces of the audio information are similar between the users; and outputting the recognition target person information and information indicating that a plurality of users having a same user group are included when the plurality of users having the same user group are included in the selected recognition targets. Park teaches calculating, from registration information in which pieces of audio information of a plurality of users are registered, group information for each user group in which features of the pieces of the audio information are similar between the users (Park, Pg. 6, paragraph 2: "The reference score setting unit 350 may determine whether the corresponding users are a general group or a special group based on the similarity of voices between pre-registered home users. That is, when the voice similarity between the users in the house exceeds a reference score (hereinafter referred to as 'first reference score' for convenience of description) for determining whether a special group is present, the reference score setting unit 350 determines the corresponding users It can be classified (recognized) as a 'special group'. On the other hand, if the voice similarity between the users in the home is less than or equal to the first reference score, the reference score setting unit 350 may classify (recognize) the users as a 'normal group'."); and outputting the recognition target person information and information indicating that a plurality of users having a same user group are included when the plurality of users having the same user group are included in the selected recognition targets (Park, Pg. 7, paragraph 6: "On the other hand, as a result of the check in step 460, when the measured similarity score is greater than the first reference score, the speaker recognition device 300 recognizes a user requesting voice registration and a pre-registered home user as a special group and is a speaker The reference score for identification may be updated (S480). At this time, the speaker recognition device 300 may update the reference score for speaker identification using the measured similarity score."; Here, the indication of a special group is seen as an output that is then used to update the reference score.). Shinto and Park are analogous arts because they each belong to the same field of speech recognition systems. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition device of Shinto to incorporate the teachings of Park to create user groups based on audio similarities. This ensures that the system can accurately identify a speaker even if other speakers have similar voices (Park, Pg. 2, paragraph 6). As such, the system works better overall, and the user experience is improved. Claim(s) 3, 7, 11, and 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Shinto, in view of Park, as applied to claims 1, 2, 5, 6, 9, 10, 13, 14, and 17-19 above, and further in view of Liu et al. (US Pat. Pub No. 2023/0386478 A1 hereinafter Liu). Regarding claim 3, the rejection of claim 1 is incorporated. Shinto, in view of Park, discloses all of the elements of the claimed invention as stated above. However, Shinto, in view of Park, fails to expressly recite wherein the processing includes calculating a value of the utterance history coefficient to be higher for a piece of the audio information having a larger number of times of recognition as an utterance in a predetermined period in past, among the registered pieces of the audio information of the users. Liu teaches wherein the processing includes calculating a value of the utterance history coefficient to be higher for a piece of the audio information having a larger number of times of recognition as an utterance in a predetermined period in past, among the registered pieces of the audio information of the users (Liu, [0254]: "Similarly, over time, application-specific words may be added or removed from speech profiles corresponding to the respective application. For example, a user (or group of users) may frequently use a specific word of phrase when requesting a function of a specific application (e.g., a request “send me to the grocery store” when using a messaging application). Frequent usage of a particular term (e.g., “send”) with a particular application may cause the term to be added as an application-specific word for the respective speech profile."). Shinto, Park, and Liu are analogous arts because they each belong to the same field of speech recognition systems. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition device of Shinto, as modified by the speaker recognition apparatus of Park, to incorporate the teachings of Liu to give a piece of repeated audio information a higher utterance history coefficient. This allows the system to change how it interacts with users based on the words or commands they often say (Liu, [0254]). As such, the system can correlate certain words or commands to certain users, and improve how it responds to user’s speech. Regarding claim 7, the rejection of claim 3 is incorporated. Shinto, in view of Park and Liu, discloses all of the elements of the claimed invention as stated above. Shinto further discloses wherein the processing includes deleting, from the recognition targets, a user whose utterance history coefficient satisfies a predetermined condition among the plurality of users having the same user group, based on the utterance history coefficients of the plurality of users having the same user group (Shinto, Pg. 10, paragraph 6: " According to the above-described processing, the voice information of the user with high priority is recognized among the received voices, and the voice information other than the user with low priority is removed."). However, Shinto fails to expressly recite when the plurality of users having the same user group are included in the recognition targets to be selected. Park further teaches when the plurality of users having the same user group are included in the recognition targets to be selected (Park, Pg. 6, paragraph 2: "The reference score setting unit 350 may determine whether the corresponding users are a general group or a special group based on the similarity of voices between pre-registered home users. That is, when the voice similarity between the users in the house exceeds a reference score (hereinafter referred to as 'first reference score' for convenience of description) for determining whether a special group is present, the reference score setting unit 350 determines the corresponding users It can be classified (recognized) as a 'special group'. On the other hand, if the voice similarity between the users in the home is less than or equal to the first reference score, the reference score setting unit 350 may classify (recognize) the users as a 'normal group'."). The same motivation for claim 1 applies equally to claim 7. Regarding claim 11, the rejection of claim 3 is incorporated. Shinto, in view of Park and Liu, discloses all of the elements of the claimed invention as stated above. Shinto further discloses wherein the processing further includes receiving, from a user of the users, selection of addition or deletion of a recognition target person to or from the recognition target person information (Shinto, Pg. 4, paragraph 4: "In this case, the registration unit 101 registers the user's voice information and a predetermined vocabulary for changing the recognition target person. The voice recognition unit 103 recognizes the voice information of the user registered in the registration unit 101 and a predetermined vocabulary among the voices received by the reception unit 102. Further, the changing unit 107 changes the recognition target person set in the setting unit 106 to the user who spoke, based on the result recognized by the voice recognition unit 103."). Regarding claim 15, the rejection of claim 3 is incorporated. Shinto, in view of Park and Liu, discloses all of the elements of the claimed invention as stated above. Shinto further discloses wherein the processing further includes: inputting a voice of a speaker; and recognizing the speaker of the input voice based on the pieces of the audio information respectively corresponding to recognition target persons of the recognition targets (Shinto, Pg. 4, paragraph 2: " In this case, the voice recognition unit 103 recognizes the voice information of the recognition target person set in the setting unit 106 among the voices received by the reception unit 102. In this configuration, even when voice information of a plurality of users is registered in the registration unit 101, it is possible to recognize only the voice of the person to be recognized by setting."). Claim(s) 4, 8, 12, and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Shinto, in view of Park, as applied to claims 1, 2, 5, 6, 9, 10, 13, 14, and 17-19 above, and further in view of Kopuri et al. (US Pat. No. 11,893,999 B1 hereinafter Kopuri). Regarding claim 4, the rejection of claim 1 is incorporated. Shinto, in view of Park, discloses all of the elements of the claimed invention as stated above. However, Shinto, in view of Park, fails to expressly recite wherein the processing includes calculating a value of the utterance history coefficient to be higher for a user corresponding to a piece of the audio information in which a day of a week or a time zone having a larger number of times of recognition of an utterance in past is closer to a current day of a week or a current time zone, among the registered pieces of the audio information of the users. Kopuri teaches wherein the processing includes calculating a value of the utterance history coefficient to be higher for a user corresponding to a piece of the audio information in which a day of a week or a time zone having a larger number of times of recognition of an utterance in past is closer to a current day of a week or a current time zone, among the registered pieces of the audio information of the users (Kopuri, Col. 14, lines 39-48: "The ML component 716 may track the behavior of various users as a factor in determining a confidence level of the identity of the user. By way of example, a user may adhere to a regular schedule such that the user is at a first location during the day (e.g., at work or at school). In this example, the ML component 716 would factor in past behavior and/or trends into determining the identity of the user that provided input to the system. Thus, the ML component 716 may use historical data and/or usage patterns over time to increase or decrease a confidence level of an identity of a user."). Shinto, Park, and Kopuri are analogous arts because they each belong to the same field of speech recognition systems. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the speech recognition device of Shinto, as modified by the speaker recognition apparatus of Park, to incorporate the teachings of Kopuri to calculate the utterance history coefficient to be higher based on recognition at certain times or days. Since a user may adhere to a regular schedule, the system can better identify a user based on their typical schedule (Kopuri, Col. 14, lines 39-48). This helps improve the accuracy of the system’s identifications. Regarding claim 8, the rejection of claim 4 is incorporated. Shinto, in view of Park and Kopuri, discloses all of the elements of the claimed invention as stated above. Shinto further discloses wherein the processing includes deleting, from the recognition targets, a user whose utterance history coefficient satisfies a predetermined condition among the plurality of users having the same user group, based on the utterance history coefficients of the plurality of users having the same user group (Shinto, Pg. 10, paragraph 6: " According to the above-described processing, the voice information of the user with high priority is recognized among the received voices, and the voice information other than the user with low priority is removed."). However, Shinto fails to expressly recite when the plurality of users having the same user group are included in the recognition targets to be selected. Park further teaches when the plurality of users having the same user group are included in the recognition targets to be selected (Park, Pg. 6, paragraph 2: "The reference score setting unit 350 may determine whether the corresponding users are a general group or a special group based on the similarity of voices between pre-registered home users. That is, when the voice similarity between the users in the house exceeds a reference score (hereinafter referred to as 'first reference score' for convenience of description) for determining whether a special group is present, the reference score setting unit 350 determines the corresponding users It can be classified (recognized) as a 'special group'. On the other hand, if the voice similarity between the users in the home is less than or equal to the first reference score, the reference score setting unit 350 may classify (recognize) the users as a 'normal group'."). The motivation for claim 1 applies equally to claim 8. Regarding claim 12, the rejection of claim 4 is incorporated. Shinto, in view of Park and Kopuri, discloses all of the elements of the claimed invention as stated above. Shinto further discloses wherein the processing further includes receiving, from a user of the users, selection of addition or deletion of a recognition target person to or from the recognition target person information (Shinto, Pg. 4, paragraph 4: "In this case, the registration unit 101 registers the user's voice information and a predetermined vocabulary for changing the recognition target person. The voice recognition unit 103 recognizes the voice information of the user registered in the registration unit 101 and a predetermined vocabulary among the voices received by the reception unit 102. Further, the changing unit 107 changes the recognition target person set in the setting unit 106 to the user who spoke, based on the result recognized by the voice recognition unit 103."). Regarding claim 16, the rejection of claim 4 is incorporated. Shinto, in view of Park and Kopuri, discloses all of the elements of the claimed invention as stated above. Shinto further discloses wherein the processing further includes: inputting a voice of a speaker; and recognizing the speaker of the input voice based on the pieces of the audio information respectively corresponding to recognition target persons of the recognition targets (Shinto, Pg. 4, paragraph 2: " In this case, the voice recognition unit 103 recognizes the voice information of the recognition target person set in the setting unit 106 among the voices received by the reception unit 102. In this configuration, even when voice information of a plurality of users is registered in the registration unit 101, it is possible to recognize only the voice of the person to be recognized by setting."). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Wu et al. (US Pat. No. 12,443,687 B1) discloses a system for user identification for touch interactions. Rosen et al. (US Pat. No. 9,728,188 B1) discloses methods and devices for ignoring similar audio being received by a system. Any inquiry concerning this communication or earlier communications from the examiner should be directed to TYLER J BECKER whose telephone number is (703)756-1271. The examiner can normally be reached M-Th, 7:15am-5:45pm PT. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Washburn can be reached at (571) 272-5551. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /TYLER BECKER/ Examiner, Art Unit 2657 /DANIEL C WASHBURN/ Supervisory Patent Examiner, Art Unit 2657
Read full office action

Prosecution Timeline

Dec 05, 2024
Application Filed
Jul 14, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12694228
REAL-TIME USER COMMUNICATION SENTIMENT DETECTION FOR DYNAMIC ANOMALY DETECTION AND MITIGATION
3y 4m to grant Granted Jul 28, 2026
Patent 12682113
SYSTEMS, METHODS, AND APPARATUSES FOR GENERATING STRUCTURED DATA FROM UNSTRUCTURED DATA USING NATURAL LANGUAGE PROCESSING TO GENERATE A SECURE MEDICAL DASHBOARD
3y 0m to grant Granted Jul 14, 2026
Patent 12651592
SYSTEM, METHOD, AND COMPUTER PROGRAM FOR REAL-TIME LANGUAGE TRANSLATION USING GENERATIVE ARTIFICIAL INTELLIGENCE
3y 0m to grant Granted Jun 09, 2026
Patent 12632657
Joint Speech and Text Streaming Model for ASR
2y 10m to grant Granted May 19, 2026
Patent 12614560
REVERBERATION REMOVAL DEVICE, PARAMETER ESTIMATION DEVICE, REVERBERATION REMOVAL METHOD, PARAMETER ESTIMATION METHOD, AND PROGRAM
2y 9m to grant Granted Apr 28, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
74%
Grant Probability
80%
With Interview (+6.3%)
2y 8m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 23 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month