Prosecution Insights
Last updated: August 06, 2026
Application No. 18/022,255

VOICE PROCESSING DEVICE FOR PROCESSING VOICE SIGNAL AND VOICE PROCESSING SYSTEM COMPRISING SAME

Non-Final OA §103
Filed
Feb 20, 2023
Priority
Aug 19, 2020 — RE 10-2020-0103909 +2 more
Examiner
SOLAIMAN, FOUZIA HYE
Art Unit
2653
Tech Center
2600 — Communications
Assignee
Amosense Co., Ltd.
OA Round
5 (Non-Final)
67%
Grant Probability
Favorable
5-6
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 67% — above average
67%
Career Allowance Rate
46 granted / 69 resolved
+4.7% vs TC avg
Strong +54% interview lift
Without
With
+54.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
7 currently pending
Career history
80
Total Applications
across all art units

Statute-Specific Performance

§101
29.3%
-10.7% vs TC avg
§103
49.8%
+9.8% vs TC avg
§102
15.6%
-24.4% vs TC avg
§112
2.3%
-37.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 69 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Response to Applicant’s Arguments and Amendments This communication is in response to continuation filed on 05/27/2026. Applicant filed an amendment on 05/27/2026., amending independent claims 1. Claims 2-4, 6, and 8-15, are cancelled. The pending claims are 1, 5, and 7. Applicant's arguments have been fully considered but they are moot because examiner used new prior art for amended claim limitation. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim 1, and 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zunn Choi, KR 101989127 B1 {IDS provided} in view of PAEK MIN HO, KR 20200012104 A {IDS provided}, and in view of by FURUKAWA et al. US 20190303443 A1 {IDS provided} and further in view of Lai et al. US 20170331782 A1 and further in view of Ekkizogloy et al. US 20180332389 A1 Regarding Claim 1, Zunn Choi teaches: [Claim 1] A voice processing device comprising: a voice data receiving circuit configured to receive input voice data associated with voices of speakers; Zunn Choi teaches (“The translator 10 includes a microphone module 11 and a speaker 12. The translator 10 determines a speaker's direction according to an embodiment of the present invention, and converts a voice signal received from the speaker's direction among input voice signals. …” page 2, 2nd. Para from bottom.) by Zunn Choi, KR 101989127 B1 {IDS provided} wherein the input voice data is generated from voice signals generated by a plurality of microphones. Zunn Choi teaches (“And the microphone module is a microphone module including a plurality of individual directional microphones.”) (“Referring to FIG. 6B, the processor 22 of the translation apparatus 10 obtains an input voice signal from the microphone module 11. In this case, the input voice signal may be a voice signal obtained from individual microphones of the microphone module 11.”) by Zunn Choi, KR 101989127 B1 {IDS provided} a voice data output circuit configured to output voice data associated with the voices of the speakers; Zunn Choi teaches (“The translator 10 includes a microphone module 11 and a speaker 12. The translator 10 determines a speaker's direction according to an embodiment of the present invention, and converts a voice signal received from the speaker's direction among input voice signals. By extracting to generate the translation target data, and to obtain the translation data for the translation target data to output the output voice signal to the speaker. The detailed operation of the translation apparatus 10 will be described later.”) page 2, 2nd. Para from bottom.) (“The translation apparatus 10 may include a microphone module 11, a speaker 12, a memory 21, a processor 22, a communication module 23, and an input / output interface 24. …” page 3, para 6) by Zunn Choi, KR 101989127 B1 {IDS provided} and a processor configured to generate a control command for outputting the output voice data, wherein the processor is further configured to: Zunn Choi teaches (“The translator 10 includes a microphone module 11 and a speaker 12. The translator 10 determines a speaker's direction according to an embodiment of the present invention, and converts a voice signal received from the speaker's direction among input voice signals. By extracting to generate the translation target data, and to obtain the translation data for the translation target data to output the output voice signal to the speaker. The detailed operation of the translation apparatus 10 will be described later.” Page 2, 2nd paragraph from bottom.) (“The processor 22 may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor 22 by the memory 21 or the communication module 23. For example, the processor 22 may be configured to execute a command received according to a program code stored in a recording device such as the memory 21.” Page 3, 3rd. paragraph from bottom.) generate plural speakers position data Zunn Choi teaches (“Abstract The present invention relates to a translation apparatus, a translation method, and a computer program for a translation method. More particularly, after determining the direction of a plurality of speakers from an input speech signal, the speech is extracted by considering each speaker's direction and translated. The ability to output a signal relates to an apparatus, a method and a program.” Page 2 first para.) (“Referring to FIG. 4, the direction determiner 311 first analyzes an input voice signal obtained from the microphone module 11 to determine a direction in which a speaker exists (S41). In this case, the input voice signal for determining the direction in which the speaker exists may be a signal of an initial setting period by the user's setting or the internal setting of the translation apparatus 10 among the input voice signals collected by the microphone module 11. More specifically, the direction determiner 311 determines a first direction in which the first speaker exists by analyzing the voice signal of the initial setting section acquired from the microphone module 11, and a second direction in which the second speaker exists. Can be determined. In this case, the first direction and the second direction may be determined using a known sound source location algorithm. The direction determiner 311 may determine the direction in which the speaker exists using the sound source location algorithm from the directivity information of the individual microphones included in the microphone module 11 and the voice signals acquired by the individual microphones. In this case, the direction in which the speaker exists may be a direction in which dB (decibels) of the input voice signal becomes a maximum value. In addition, the first direction and the second direction may be determined based on the translation apparatus 10.”) (“Next, the speaker determiner 312 determines whether the input direction of the input voice signal is within the first direction and the error range or within the second direction and the error range to determine the speaker of the input voice signal. At this time, the input direction is within the error range with the first direction or the second direction, the angle (for example, the direction derived from the sound source location algorithm) of the current input voice signal is a predetermined angle from the first direction or the second direction. It can mean mine. That is, the speaker determiner 312 of the present invention determines whether the input voice signal currently collected and recorded by the first speaker is based on the first direction which is the direction of the first speaker and the second direction which is the direction of the second speaker. Or by the second speaker.”) by Zunn Choi, KR 101989127 B1 {IDS provided} Fig. 4-5, Zunn Choi, teaches (“At this time, the input direction is within the error range with the first direction or the second direction, the angle (for example, the direction derived from the sound source location algorithm) of the current input voice signal is a predetermined angle from the first direction or the second direction. It can mean mine. That is, the speaker determiner 312 of the present invention determines whether the input voice signal currently collected and recorded by the first speaker is based on the first direction which is the direction of the first speaker and the second direction which is the direction of the second speaker. Or by the second speaker.” Page 6, last para and page 7, first para) (“… Similarly, when the speaker determiner 312 determines that the speaker of the input voice signal is the second speaker, the translation target data generator 313 extracts the voice signal received in the second direction based on the input voice signal, The second translation target data is generated from the voice signal received in the second direction.” Page 7, 2nd para.) by Zunn Choi, KR 101989127 B1 {IDS provided} first output voice data associated with a voice of the first speaker by using the input voice data, Zunn Choi teaches (“… . It can mean mine. That is, the speaker determiner 312 of the present invention determines whether the input voice signal currently collected and recorded by the first speaker is based on the first direction which is the direction of the first speaker and the second direction which is the direction of the second speaker. Or by the second speaker.”) (“Next, the translation data acquisition unit 314 may obtain the translation data for the translation target data and output the output voice signal to the speaker. In more detail, the translation data acquisition unit 314 obtains the first translation data for the first translation target data and outputs an output speech signal to the speaker, or obtains the first translation data for the second translation target data and outputs the output speech. The signal can be output to the speaker. In this case, the first translation data may be a translation of the first translation target data into a language of the second translation target data, and the second translation data may be a translation of the second translation target data into a language of the first translation target data.”) by Zunn Choi, KR 101989127 B1 {IDS provided} second output voice data associated with a voice of the second speaker by using the input voice data, Zunn Choi teaches (“… . It can mean mine. That is, the speaker determiner 312 of the present invention determines whether the input voice signal currently collected and recorded by the first speaker is based on the first direction which is the direction of the first speaker and the second direction which is the direction of the second speaker. Or by the second speaker.”) (“Next, the translation data acquisition unit 314 may obtain the translation data for the translation target data and output the output voice signal to the speaker. In more detail, the translation data acquisition unit 314 obtains the first translation data for the first translation target data and outputs an output speech signal to the speaker, or obtains the first translation data for the second translation target data and outputs the output speech. The signal can be output to the speaker. In this case, the first translation data may be a translation of the first translation target data into a language of the second translation target data, and the second translation data may be a translation of the second translation target data into a language of the first translation target data.”) by Zunn Choi, KR 101989127 B1 {IDS provided} Zunn Choi further teaches: determine a first source language data matched and stored with the first position data among the source language data, Zunn Choi teaches the memory 21 may store program codes and settings for controlling the translation apparatus 10 and input voice signals temporarily or permanently.” Page 3, para 4 from bottom page) (“… the speaker determiner 312 of the present invention determines whether the input voice signal currently collected and recorded by the first speaker is based on the first direction which is the direction of the first speaker and the second direction which is the direction of the second speaker. Or by the second speaker. (“Next, the translation target data generation unit 313 extracts the voice signal received in the direction in which the speaker exists (first direction or second direction) based on the input voice signal, and from the voice signal received in the direction. The target data for translation is generated (S42). At this time, when the speaker determiner 312 determines that the speaker of the input voice signal is the first speaker, the translation target data generator 313 extracts the voice signal received in the first direction based on the input voice signal, The first translation target data is generated from the voice signal received in the first direction. Similarly, when the speaker determiner 312 determines that the speaker of the input voice signal is the second speaker, the translation target data generator 313 extracts the voice signal received in the second direction based on the input voice signal, The second translation target data is generated from the voice signal received in the second direction.” Page 7, 2nd. paragraph) (“In addition, the translation target data generated by the translation target data generation unit 313 may include not only a voice signal but also speaker information corresponding to the voice signal, direction information of the speaker, or speaker information. Accordingly, the user terminal 110 or the server 150 may determine the language and the translation language of the translation target data based on the speaker information or the direction information without having to recognize the language of the translation target data each time the acquisition target data is acquired. Translation time can be reduced.” Page 7, 3rd. paragraph) (“The translation apparatus 10 may include a microphone module 11, a speaker 12, a memory 21, a processor 22, a communication module 23, and an input / output interface 24. The memory 21 is a computer-readable recording medium, and may include a permanent mass storage device such as random access memory (RAM), read only memory (ROM), and a disk drive. In addition, the memory 21 may store program codes and settings for controlling the translation apparatus 10 and input voice signals temporarily or permanently.” Page 3, 3rd para from bottom.) (“In this case, the first direction and the second direction may be determined using a known sound source location algorithm. The direction determiner 311 may determine the direction in which the speaker exists using the sound source location algorithm from the directivity information of the individual microphones included in the microphone module 11 and the voice signals acquired by the individual microphones. In this case, the direction in which the speaker exists may be a direction in which dB (decibels) of the input voice signal becomes a maximum value. In addition, the first direction and the second direction may be determined based on the translation apparatus 10. …” page 6, 2nd para, line 13-6, from bottom page ) (“Next, the speaker determiner 312 determines whether the input direction of the input voice signal is within the first direction and the error range or within the second direction and the error range to determine the speaker of the input voice signal. At this time, the input direction is within the error range with the first direction or the second direction, the angle (for example, the direction derived from the sound source location algorithm) of the current input voice signal is a predetermined angle from the first direction or the second direction. It can mean mine. That is, the speaker determiner 312 of the present invention determines whether the input voice signal currently collected and recorded by the first speaker is based on the first direction which is the direction of the first speaker and the second direction which is the direction of the second speaker. Or by the second speaker.” Page 6, last paragraph) by Zunn Choi, KR 101989127 B1 {IDS provided} Zunn Choi does not explicitly teach determine a second source language data matched and stored with the second position data among the source language data. PAEK MIN HO teaches: transmit, to the voice data output circuit, the control command for outputting the first output voice data to a translation environment for translating first source language into the target language corresponding the target language data PAEK MIN HO teaches (“The present invention relates to a real-time multi-interpretation wireless earset and method, and more particularly, the real-time multi-interpretation wireless earset transmits and receives the data combined with the ID signal for the corresponding language information through the translation server translation in real time”) (“That is, the second earset 112 transmits the identification ID (language) and voice data received from the first earset 111 to the translation server 200, and the TTS which is a value processed by the translation server 200. The data is received and output as voice data.”) (“ID definition step (S340) is a step in which the translation server 200 defines the ID, after determining the communication method by identifying the earset 100 as in the previous step, and receives the data transmitted from the earset 100 According to the ID data stored in each earset 100, the data is generated in each language to prepare for transmission. Therefore, when a signal carrying voice data is transmitted from any one of the earsets 100, the language-specific data of all the ID data determined in the ID definition step S340 is generated and transmitted to the earset 100 in real time.”) by PAEK MIN HO, KR 20200012104 PAEK MIN HO is considered to be analogous to the claimed invention because it relates to a real-time multi-interpretation wireless earset and method, and more particularly, the real-time multi-interpretation wireless earset transmits and receives the data combined with the ID signal for the corresponding language information through the translation server translation in real time. Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Zunn Choi to incorporate the teachings of PAEK MIN HO in order to add feature determine a first and second source language data matched and stored with the first position data among the source language data. One could have been motivated to do so because the system can conduct useful conversation. (“… here is a problem that the cost of text to use. In addition, although using the smart phone to use the interpretation service, this also is useful for the conversation through meetings, ….”) by PAEK MIN HO, KR 20200012104 A The combination does not explicitly teach determine a second source language data matched and stored with the second position data among the source language data. FURUKAWA teaches: determine first position data corresponding to the first speaker position data among the stored predefined position data, FURUKAWA teaches using the positional relationship indicated in one layout information selected in advance among the plurality of layout information stored in Based on the sound source direction. (“[0052] …using a positional relationship indicated by a layout information item selected in advance from a plurality of layout information items that are stored in storage and respectively indicate different positional relationships between the user, … … the first area corresponding to a position of the identified one of the user and the conversation partner, the second area corresponding to a position of the other one of the user and the conversation partner.”) (“[0093] In this embodiment, storage 11 stores a plurality of layout information items respectively indicating different positional relationships of user 51, conversation partner 52, and display 300. In Storage 1, one layout information item is selected in advance from among the plurality of layout information items that are stored therein.”) (“[0094] In addition, storage 11 stores a coordinate system centering speech translation apparatus 100 and indices assigned respectively to segment areas of a region centering speech translation apparatus 100.”) (“[0052] … (i) identifies that an utterer who utters speech is one of the user and the conversation partner, based on the sound source direction estimated by the sound source direction estimator after the start of the translation is instructed by the translation start button, using a positional relationship indicated by a layout information item selected in advance from a plurality of layout information items that are stored in storage and respectively indicate different positional relationships between the user, the conversation partner, and a display, and (ii) determines a translation direction indicating an input language in which content of the acoustic signal is recognized and an output language into which the content of the acoustic signal is translated, the input language being one of a first language and a second language and the output language being the other one of the first language and the second language; a translator which obtains, according to the translation direction determined by the controller, (i) original text indicating the content of the acoustic signal obtained by causing a recognition processor to recognize the acoustic signal in the input language and (ii) translated text indicating the content of the acoustic signal obtained by causing a translation processor to translate the original text into the output language; and a display unit which displays the original text on a first area of the display, and displays the translated text on a second area of the display, the first area corresponding to a position of the identified one of the user and the conversation partner, the second area corresponding to a position of the other one of the user and the conversation partner.”) (“[0073] … using a positional relationship indicated by a layout information item selected in advance from a plurality of layout information items that are stored in storage and respectively indicate different positional relationships between the user, the conversation partner, and a display, and (ii) determining a translation direction indicating an input language in which content of the acoustic signal is recognized and an output language into which the content of the acoustic signal is translated,. ”) by FURUKAWA et al. US 20190303443 A1 determine second position data corresponding to the second speaker position data the among stored predefined position data, FURUKAWA teaches using a positional relationship indicated by a layout information item selected in advance from a plurality of layout information items that are stored in storage and respectively indicate different positional relationships between the user, the conversation partner. (“[0066] In addition, for example, the speech translation apparatus may further include: a speech determiner which determines whether the acoustic signal obtained by the microphone array unit includes speech, wherein the controller may determine the translation direction only when (i) the acoustic signal is determined to include speech by the speech determiner and (ii) the sound source direction estimated by the sound source direction estimator indicates the position of the user or the position of the conversation partner in the positional relationship indicated by the layout information item.”) (“[0073] In addition, a recording medium according to the present disclosure is a non-transitory computer-readable recording medium having a program stored thereon for causing a speech translation apparatus to execute a speech translation method, the speech translation apparatus including a translation start button which instructs start of translation when operated by one of a user of the speech translation apparatus and a conversation partner of the user, the speech translation method including: estimating a sound source direction by processing an acoustic signal obtained by a microphone array unit; (i) identifying that an utterer who utters speech is one of the user and the conversation partner of the user, based on the sound source direction estimated by the sound source direction estimator after the start of the translation is instructed by the translation start button, using a positional relationship indicated by a layout information item selected in advance from a plurality of layout information items that are stored in storage and respectively indicate different positional relationships between the user, the conversation partner, and a display, and (ii) determining a translation direction indicating an input language in which content of the acoustic signal is recognized and an output language into which the content of the acoustic signal is translated, the input language being one of a first language and a second language and the output language being the other one of the first language and the second language; obtaining, according to the translation direction determined in the determining, (i) original text indicating the content of the acoustic signal obtained by causing a recognition processor to recognize the acoustic signal in the input language and (ii) translated text indicating the content of the acoustic signal obtained by causing a translation processor to translate the original text into the output language; and displaying the original text on a first area of the display, and displaying the translated text on a second area of the display, the first area corresponding to a position of the identified one of the user and the conversation partner, the second area corresponding to a position of the other one of the user and the conversation partner.”) by FURUKAWA et al. US 20190303443 A1 FURUKAWA further teaches: determine a second source language data matched and stored with the second position data among the source language data FURUKAWA teaches using the positional relationship indicated in one layout information selected in advance among the plurality of layout information stored in Based on the sound source direction. (“[0052] …using a positional relationship indicated by a layout information item selected in advance from a plurality of layout information items that are stored in storage and respectively indicate different positional relationships between the user, … … the first area corresponding to a position of the identified one of the user and the conversation partner, the second area corresponding to a position of the other one of the user and the conversation partner.”) (“[0093] In this embodiment, storage 11 stores a plurality of layout information items respectively indicating different positional relationships of user 51, conversation partner 52, and display 300. In Storage 1, one layout information item is selected in advance from among the plurality of layout information items that are stored therein.”) (“[0094] In addition, storage 11 stores a coordinate system centering speech translation apparatus 100 and indices assigned respectively to segment areas of a region centering speech translation apparatus 100.”) (“[0096] The layout information item illustrated in FIG. 4A indicates a positional relationship in the case where speech translation apparatus 100 is used in portrait orientation by user 51 and conversation partner 52 facing each other. More specifically, the layout information item indicates the positional relationship in which user 51 who speaks the first language is present at the bottom side with respect to center line L.sub.1 which divides display 300 into top and bottom areas, conversation partner 52 who speaks the second language is present at the top side with respect to center line L.sub.1, and user 51 and conversation partner 52 face each other. Alternatively, the layout information item illustrated in FIG. 4A may indicate a positional relationship in which user 51 who speaks the first language is present in sound source direction 61 of the bottom side of speech translation apparatus 100 used in portrait orientation and conversation partner 52 who speaks the second language is present in sound source direction 62 of the top side of speech translation apparatus 100. In this way, FIG. 4A illustrates the layout information item indicating the positional relationship in which user 51 and conversation partner 52 face each other across display 300.”) (“[0118] More specifically, when user 51 is identified as the utterer, controller 13 determines the translation direction specifying the input language in which the content of the acoustic signal is recognized as the first language and the output language into which the content of the acoustic signal is to be translated as the second language. It is to be noted that controller 13 may determine a translation direction from the first language to the second language when user 51 is identified as the utterer. When conversation partner is identified as the utterer, controller 13 determines a translation direction specifying the input language as the second language and the output language as the first language. Controller 13 controls translator 14 according to the determined translation direction. It is to be noted that controller 13 may determine a translation direction from the second language to the first language when conversation partner 52 is identified as the utterer.”) voice source position of a first and second speaker (“[0066] In addition, for example, the speech translation apparatus may further include: a speech determiner which determines whether the acoustic signal obtained by the microphone array unit includes speech, wherein the controller may determine the translation direction only when (i) the acoustic signal is determined to include speech by the speech determiner and (ii) the sound source direction estimated by the sound source direction estimator indicates the position of the user or the position of the conversation partner in the positional relationship indicated by the layout information item.”) by FURUKAWA et al. US 20190303443 A1 FURUKAWA is considered to be analogous to the claimed invention because it relates to speech translation apparatus. Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Zunn Choi and PAEK MIN HO to further incorporate the teachings of FURUKAWA in order to add feature determine a first and second source language data matched and stored with the first position data among the source language data. One could have been motivated to do so because the system can have accurate translation recognition. (“[0202] … More specifically, speech translation apparatus 100B according to this variation can use an acoustic signal representing only the utterance(s) of user 51 or conversation partner 52, it is possible to increase the recognition accuracy and the translation accuracy of the acoustic signal.”) by FURUKAWA et al. US 20190303443 A1 The combination does not explicitly teach determine a target language data of the first speaker as the second source language data matched and stored with the second position data. The combination does not explicitly teach memory configured to store Predefined position data and source language data by matching them with each other. Lai teaches: memory configured to store Predefined position data and source language data by matching them with each other Lai teaches (“[0045] For example, in the context of the current invention, a collection of data records in data storage may comprise a user profile, associated with the user id, which define the user's native language or location, from which the user's native language may be extrapolated. The user's language and location may be determined from, for example, a web page or browser setting allowing the user to set the user's language or location, an IP address defining the user's location, a Top-Level Domains (TLD) that implies a user's preferred language, etc.”) (“[0093] Using the identified secondary language, server(s) 110 may replace tokens in the second level domain (SLD) of the candidate domain name with tokens in the secondary language. In some embodiments, tokens may be replaced by the closest translation of tokens in the secondary language, according to the translation relationships 225 within the language map 220. In some embodiments, server(s) 110 may determine the tokens to replace existing tokens in the SLD according to the most frequently used tokens identified from the stored zone data 235. The translated list of candidate domain names may then be transmitted to the client computer 120.”) (“[0094] In some embodiments, server(s) 110 may concatenate one or more additional tokens to the tokens in the domain name candidates, based on the adjacent token data in the stored zone 235 and usage 240 data. In embodiments where multiple domain names are available in the user's primary language for the user's selected Top-Level Domains (TLD) the candidate domain names may include the most frequent adjacent tokens in that language according to the stored zone 235 and usage 240 data. In embodiments where translated tokens in one or more languages are required (e.g., lack of sufficient candidate domains for TLD, user co-mingles languages in requested domain, etc.), the adjacent tokens may be concatenated in the primary language or may be translated into the most frequently used language for translation according to the stored zone 235 and usage 240 data.”) by Lai et al. US 20170331782 A1 determine a target language data of the first speaker as the second source language data matched and stored with the second position data Lai teaches (“[0045] For example, in the context of the current invention, a collection of data records in data storage may comprise a user profile, associated with the user id, which define the user's native language or location, from which the user's native language may be extrapolated. The user's language and location may be determined from, for example, a web page or browser setting allowing the user to set the user's language or location, an IP address defining the user's location, a Top-Level Domains (TLD) that implies a user's preferred language, etc.”) (“[0093] Using the identified secondary language, server(s) 110 may replace tokens in the second level domain (SLD) of the candidate domain name with tokens in the secondary language. In some embodiments, tokens may be replaced by the closest translation of tokens in the secondary language, according to the translation relationships 225 within the language map 220. In some embodiments, server(s) 110 may determine the tokens to replace existing tokens in the SLD according to the most frequently used tokens identified from the stored zone data 235. The translated list of candidate domain names may then be transmitted to the client computer 120.”) (“[0094] In some embodiments, server(s) 110 may concatenate one or more additional tokens to the tokens in the domain name candidates, based on the adjacent token data in the stored zone 235 and usage 240 data. In embodiments where multiple domain names are available in the user's primary language for the user's selected Top-Level Domains (TLD) the candidate domain names may include the most frequent adjacent tokens in that language according to the stored zone 235 and usage 240 data. In embodiments where translated tokens in one or more languages are required (e.g., lack of sufficient candidate domains for TLD, user co-mingles languages in requested domain, etc.), the adjacent tokens may be concatenated in the primary language or may be translated into the most frequently used language for translation according to the stored zone 235 and usage 240 data.”) by Lai et al. US 20170331782 A1 Lai is considered to be analogous to the claimed invention because it relates to the field of Domain Name registration and specifically to the field of generating, and validating the characters and tokens within, a list of suggested candidate domain names. Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Zunn Choi, PAEK MIN HO and FURUKAWA to further incorporate the teachings of Lai in order to add feature determine a first and second source language data matched and stored with the first position data among the source language data. One could have been motivated to do so because the system can improve the efficiency of the presentation of these additional candidate domain names, the registrar's servers may detect the characters entered by a user into a domain name request user interface control. (“[0032] Further, to improve the efficiency of the presentation of these additional candidate domain names, the registrar's servers may detect the characters entered by a user into a domain name request user interface control (e.g., the domain name search text box/drop down seen in FIG. 8, described in more detail below), generate the candidate domain names in real time, and modify the user interface control to display the candidate domain names. The display may also include additional user interface controls to display the candidate domain names using combinations of additional languages.”) by Lai et al. US 20170331782 A1 The combination does not teach generate first/ second speaker position data representing a position of a first/second speaker among the speakers based on a distance between the plurality of microphones and times when the voice signals are received by the plurality of microphones. Ekkizogloy teaches: generate first speaker position data representing a voice source position of a first speaker among the speakers based on a distance between the plurality of microphones and times when the voice signals are received by the plurality of microphones. Ekkizogloy teaches (“[0029] In a multi-passenger vehicle there may be multiple sources of audio at any one time.”) (“[0032] FIG. 2 shows a simplified diagram 200 of an array of microphones M1-M3 placed around an audio source 210, according to certain embodiments. Using multiple microphones disposed at different positions from an audio source can both improve the fidelity of the recording and provide audio location capabilities using audio phase and/or timing analysis, as further discussed below. Referring to FIG. 2, audio source 210 emits an audio signal 220, which is picked up by microphones M1-M3. Microphone M1 is at a distance L1 from audio source 210, microphone M2 is at a distance L2 from audio source 210, and microphone M3 is at a distance L3 from audio source 210. Each microphone M1-M3 can be disposed at a different position relative to audio source 210. As shown in the example portrayed in FIG. 2, M1 is the closest and M2 is the farthest way from audio source 210. Each microphone M1-M3 may receive audio signal 220 (i.e., audio data) at a different time depending on their relative position with respect to audio source 210. These time differences can be calculated (i.e., as phase differences), and used to determine the location of audio source 210, such as by trilateration, as would be understood by one of ordinary skill in the art.”) (“[0034] FIG. 3 shows a graph 300 of audio recordings for a multi-microphone array, according to certain embodiments. Returning to a simple three-microphone example, graph 300 depicts amplitude vs. time for the audio data received by each of microphones M1-M3, as shown in FIG. 2. Microphone M1 receives audio signal 220 at time t1, microphone M2 receives audio signal 220 at time t2, and microphone M3 receives audio signal 220 at time t3. As mentioned above, the time deltas between the received signals (e.g., ΔM1-M2, ΔM1-M3, ΔM2-M3) can be used to determine a location of audio source 210 using conventional trilateration techniques including time difference of arrival (TDOA), cross-correlation functions between audio signals, and geometric principles, as would be understood by one of ordinary skill in the art.”) (“(“[0054] At step 840, method 800 can include isolating second audio data received from the plurality of microphones, taking into account the determined location of the source. The second audio data can be isolated in a number of ways. For instance, the plurality of microphones can be physically adjusted to move their directional focus to the location of the source of the first audio data. …”) by Ekkizogloy et al. US 20180332389 A1 Ekkizogloy further teaches: generate second speaker position data representing a voice source position of a second speaker among the speakers based on the distance between the plurality of microphones and the times when the voice signals are received by the plurality of microphones, Ekkizogloy teaches (“[0029] In a multi-passenger vehicle there may be multiple sources of audio at any one time.”) (“[0032] FIG. 2 shows a simplified diagram 200 of an array of microphones M1-M3 placed around an audio source 210, according to certain embodiments. Using multiple microphones disposed at different positions from an audio source can both improve the fidelity of the recording and provide audio location capabilities using audio phase and/or timing analysis, as further discussed below. Referring to FIG. 2, audio source 210 emits an audio signal 220, which is picked up by microphones M1-M3. Microphone M1 is at a distance L1 from audio source 210, microphone M2 is at a distance L2 from audio source 210, and microphone M3 is at a distance L3 from audio source 210. Each microphone M1-M3 can be disposed at a different position relative to audio source 210. As shown in the example portrayed in FIG. 2, M1 is the closest and M2 is the farthest way from audio source 210. Each microphone M1-M3 may receive audio signal 220 (i.e., audio data) at a different time depending on their relative position with respect to audio source 210. These time differences can be calculated (i.e., as phase differences), and used to determine the location of audio source 210, such as by trilateration, as would be understood by one of ordinary skill in the art.”) (“[0034] FIG. 3 shows a graph 300 of audio recordings for a multi-microphone array, according to certain embodiments. Returning to a simple three-microphone example, graph 300 depicts amplitude vs. time for the audio data received by each of microphones M1-M3, as shown in FIG. 2. Microphone M1 receives audio signal 220 at time t1, microphone M2 receives audio signal 220 at time t2, and microphone M3 receives audio signal 220 at time t3. As mentioned above, the time deltas between the received signals (e.g., ΔM1-M2, ΔM1-M3, ΔM2-M3) can be used to determine a location of audio source 210 using conventional trilateration techniques including time difference of arrival (TDOA), cross-correlation functions between audio signals, and geometric principles, as would be understood by one of ordinary skill in the art.”) (“[0054] At step 840, method 800 can include isolating second audio data received from the plurality of microphones, taking into account the determined location of the source. The second audio data can be isolated in a number of ways. For instance, the plurality of microphones can be physically adjusted to move their directional focus to the location of the source of the first audio data. …”) by Ekkizogloy et al. US 20180332389 A1 Ekkizogloy is considered to be analogous to the claimed invention because it relates to vehicular systems, and in particular to systems and methods to detect and isolate audio in a vehicle using multiple microphones. Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Zunn Choi and PAEK MIN HO and FURUKAWA and Lai to incorporate the teachings of Ekkizogloy in order to add the localization module which is configure to determine user location of the user relative to microphones. One could have been motivated to do so because the system can improve audio reception. (“[0032] … Using multiple microphones disposed at different positions from an audio source can both improve the fidelity of the recording and provide audio location capabilities using audio phase and/or timing analysis,...”) by Ekkizogloy et al. US 20180332389 A1 Regarding Claim 7, the combination teaches the device claim 1 as identified above. Zunn Choi further teaches: 7. (Previously Presented) The voice processing device of claim 1 wherein the processor is configured to: generate second output voice data associated with a voice of the second speaker by using the input voice data, and Zunn Choi, teaches (“That is, the translation apparatus 10 of the present invention may translate and output the voice of the first speaker into the language of the second speaker, and may translate and output the voice of the second speaker into the language of the first speaker. At this time, the voice of each speaker is automatically extracted in consideration of the direction, so that the translation device 10 does not need to exist near the speaker who is currently speaking like the existing microphone, so that the translation device 10 is located near the first and second speakers. The voice of the first speaker can be translated into the voice of the language of the second speaker and vice versa just by placing it therein.” page 7, last para) transmit, to the voice data output circuit, the control command for outputting the first output voice data to a translation environment for translating the second source language into the first source language. Zunn Choi, teaches (“ The language recognition unit 321 recognizes a language of the translation target data based on the translation target data obtained from the translation apparatus 10. The character conversion unit 322 converts the translation target data, which is voice data, into characters based on the recognized language or based on a language corresponding to the speaker's direction information or speaker information. The translation unit 323 translates the translation target data converted into text into another language. The voice converter converts the text data translated into another language into voice data, generates the translation target data, and transmits the translation target data to the translation apparatus 10. For example, an application of the user terminal 110 may obtain translation target data from the translation apparatus 10 and transmit the translation target data to the server 150. The processor 222 of the server 150 recognizes the language of the data to be translated, converts the data to be translated into text data based on the recognized language, translates the translated text data into another language, and then generates translated data. The translation data may be transmitted to the user terminal 110. The user terminal 110 may transmit the obtained translation data to the translation apparatus 10.” page 6, paragraph 3.) (“The translated Korean is delivered to a user who speaks Korean, and the user hears the contents and answers the language in Korean. The Korean-speaking user's pendant recognizes the Korean language and delivers Korean to the English-speaking user's pendant. The pendant is translated into English and output to the user.” page 4, Lines 25-28) (“That is, the translation apparatus 10 of the present invention may translate and output the voice of the first speaker into the language of the second speaker, and may translate and output the voice of the second speaker into the language of the first speaker. At this time, the voice of each speaker is automatically extracted in consideration of the direction, so that the translation device 10 does not need to exist near the speaker who is currently speaking like the existing microphone, so that the translation device 10 is located near the first and second speakers. The voice of the first speaker can be translated into the voice of the language of the second speaker and vice versa just by placing it therein.” page 7, last para) (“Referring to FIG. 5A, the translation apparatus 10 and the user terminal 110 are placed on the table 53, and the first speaker 51 is positioned in the first direction 51 ′ based on the translation apparatus 10. The second speaker 52 may be located in the second direction 52 ′. According to the present invention, the conversation between the speakers is translated with the translator 10 within 5m distance, preferably 1 to 2m distance from the speaker, without having to take the microphone near the speaker when collecting the speaker's voice. can do. The translator 10 extracts a voice signal in the first direction 51 ', converts it into a language of the second speaker 52, outputs it to the speaker 12, and extracts a voice signal in the second direction 52'. Can be converted into the language of the first speaker 51 and output to the speaker 12. In addition, a signal for controlling the translation apparatus 10 may be input with the user terminal 110 within a predetermined distance from the translation apparatus 10, or information indicating the current state of the translation apparatus 10 may be displayed.” Page 2nd. para) by Zunn Choi, KR 101989127 B1 {IDS provided} Claim 5 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zunn Choi, PAEK MIN HO, FURUKAWA, Lai and Ekkizogloy in view of NAKADAI et al. US 20150154957 A1. Regarding Claim 5, the combination teaches the device claim 1 as identified above. Zunn Choi further teaches: 5. (Original) The voice processing device of claim 1, wherein the processor is configured to convert the first output voice data associated with the voice of the first speaker into text data that is expressed in the first source language, and Zunn Choi teaches (“The existing translation apparatus processes the voice of a user input through a microphone and outputs the translated text or voice. …” page 2, 2nd. Para.)(“The language recognition unit 321 recognizes a language of the translation target data based on the translation target data obtained from the translation apparatus 10. The character conversion unit 322 converts the translation target data, which is voice data, into characters based on the recognized language or based on a language corresponding to the speaker's direction information or speaker information. The translation unit 323 translates the translation target data converted into text into another language.” Page 6, 3rd paragraph.) (“…For example, when the user terminal 1 110 transmits a text data conversion and translation request of the translation target data to the server 150 through a network under the control of the application, the server 150 converts the translation target data into the text data. After the translation, the translation data may be transmitted to the user terminal 1 110, and the user terminal 1 110 may provide the translation device 10 under the control of the application. …” page 3, lines 11-16) by Zunn Choi, KR 101989127 B1 PAEK MIN HO teaches: wherein the voice data output circuit is configured to transmit the text data converted under the control of the processor to the translation environment. PAEK MIN HO teaches (“The translation server 200 analyzes the received identification ID and voice data to perform language translation. The translated language is to be converted to TTS. The data converted into TTS by the translation server 200 is transferred to the second earset 112, and the second earset 112 outputs the received TTS data as voice data. The translation server 200 may be a method of communicating with the ear set by having a separate server.” Page 3, 4th paragraph from bottom page) by PAEK MIN HO, KR 20200012104 A PAEK MIN HO is considered to be analogous to the claimed invention because it relates to a real-time multi-interpretation wireless earset and method, and more particularly, the real-time multi-interpretation wireless earset transmits and receives the data combined with the ID signal for the corresponding language information through the translation server translation in real time. Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Zunn Choi to incorporate the teachings of PAEK MIN HO in order to add feature determine a first and second source language data matched and stored with the first position data among the source language data. One could have been motivated to do so because the system can conduct useful conversation. (“… here is a problem that the cost of text to use. In addition, although using the smart phone to use the interpretation service, this also is useful for the conversation through meetings, ….”) by PAEK MIN HO, KR 20200012104 A The combination does not explicitly teach transmission of text data conversion. NAKADAI teaches: transmission of text data conversion FIG. 10, NAKADAI teaches (“[0005] … The transmitter includes a microphone, a speech recognition circuit, and a transmitter unit and the transmitter unit transmits character information corresponding to the details recognized on the basis of the speech recognition result to the receiver. The receiver includes a receiver unit, a central processing unit (CPU), and a display unit and displays characters on the display unit when character information is received from the transmitter.”) (“[0175] This embodiment describes the example where the plurality of speakers use the conversation support apparatus 1A, but the invention is not limited to this example. The conversation support apparatus 1A may be used by a single speaker. For example, when the speaker registers Japanese as the language in the initial state and utters speech in English, the conversation support apparatus 1A may translate the English speech uttered by the speaker into Japanese which is the registered language and may display the Japanese text in the presentation area of the image corresponding to the speaker. Accordingly, the conversation support apparatus 1A according to this embodiment can achieve an effect of foreign language learning support.”) (“[0101] (Step S5) The display image generating unit 142 generates character data corresponding to the recognition data for each speaker input from the speech recognizing unit 13 and outputs the generated character data for each speak to the image display unit 15. The display image generating unit 142 generates an image of the information indicating the direction of each speaker on the basis of the information indicating the directions of the speakers, which is input from the speech recognizing unit 13, and outputs the image of the generated information indicating the direction of each speaker to the image synthesizing unit 143.”) (“[0104] An example of the result of an experiment which is performed using the conversation support apparatus 1 according to this embodiment will be described below. FIG. 10 is a diagram showing an experiment environment.”) (“[0140] The translation unit 24 translates the speech details if necessary on the basis of the speech details, the information indicating the speakers, and the information indicating a language for each speaker which are input from the speech recognizing unit 13A, adds or replaces information indicating the translated speech details to or for the information input from the speech recognizing unit 13A, and outputs the resultant to the image processing unit 14. Specifically, an example where two speakers of the first speaker Sp1 and the second speaker Sp2 are present as the speakers, the language of the first speaker Sp1 is Japanese, the language of the second speaker Sp2 is English will be described below with reference to FIG. 14. In this case, the translation unit 24 translates the speech details so that the images 534A to 534D displayed in the second character presentation image 532 are translated from Japanese in which the first speaker Sp1 utters speech to English which is the language of the second speaker Sp2 and are then displayed. The translation unit 24 translates the speech details so that the images 524A to 524C displayed in the first character presentation image 522 are translated from English in which the second speaker Sp2 utters speech to Japanese which is the language of the first speaker Sp1 and are then displayed.”) by NAKADAI et al. US 20150154957 A1 NAKADAI is considered to be analogous to the claimed invention because it relates to a speech translation apparatus, a speech translation method, and a recording medium. Therefore, it would have been obvious for someone of ordinary skill in the art before the effective filing date of the claimed invention to modify Zunn Choi and PAEK MIN HO, FURUKAWA, Lai and Ekkizogloy to incorporate the teachings of NAKADAI in order to add a translation environment in the system. One could have been motivated to do so because the system can display the conversation details on the screen can be maintained and thus the convenience for the speakers is improved. (“[0095] … Accordingly, even when the positions of the speakers are interchanged, the conversation details displayed on the screen can be maintained and thus the convenience for the speakers is improved.”) by NAKADAI et al. US 20150154957 A1 Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to FOUZIA HYE SOLAIMAN whose telephone number is (571)270-5656. The examiner can normally be reached M-F (8-5)AM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached at (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /F.H.S./ Examiner, Art Unit 2653 /Paras D Shah/Supervisory Patent Examiner, Art Unit 2653 06/29/2026
Read full office action

Prosecution Timeline

Show 4 earlier events
Dec 19, 2025
Request for Continued Examination
Jan 13, 2026
Response after Non-Final Action
Feb 05, 2026
Non-Final Rejection mailed — §103
Apr 03, 2026
Response Filed
May 01, 2026
Final Rejection mailed — §103
May 27, 2026
Request for Continued Examination
May 29, 2026
Response after Non-Final Action
Jul 02, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12676137
System and Method for Multi-Channel Speech Privacy Processing
3y 1m to grant Granted Jul 07, 2026
Patent 12639528
LARGE LANGUAGE MODEL AND DETERMINISTIC CALCULATOR SYSTEMS AND METHODS
2y 9m to grant Granted May 26, 2026
Patent 12626066
EXTRACTING CONVERSATIONAL RELATIONSHIPS BASED ON SPEAKER PREDICTION AND TRIGGER WORD PREDICTION
3y 1m to grant Granted May 12, 2026
Patent 12592217
SYSTEM AND METHOD FOR SPEECH PROCESSING
3y 1m to grant Granted Mar 31, 2026
Patent 12579976
USER TERMINAL, DIALOGUE MANAGEMENT SYSTEM, CONTROL METHOD OF USER TERMINAL, AND DIALOGUE MANAGEMENT METHOD
3y 0m to grant Granted Mar 17, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
67%
Grant Probability
99%
With Interview (+54.0%)
2y 11m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 69 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month