DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 6/3/2026 has been entered.
Response to Amendment
In response to the office action from 3/3/2026, the applicant representative has filed a request for continued examination, filed 6/3/2026, amending claims 1-2, 4, 6-10, 12-15, and 17-20, cancelling claims 3 and 11, while arguing to traverse Alice 101 rejections. Unfortunately the latest amendments had cancelled what was determined prior art allowable subject matter, therefore the claims are now rejected in view of Weber (US 2019/0171716) and further in view of Kyu et al. (KR102045761) and for the reasons explained in the response to arguments.
Response to Arguments
Page 7, the first paragraph provides a broad overview of the latest amendments.
Page 7 section “Step 2A(1)” simply provides a copy of claim 1 and merely concludes: “These steps are not mental processes; rather, it is” done by a “computing device”; next in the section titled “Step 2A(2)” on page 7, it is concluded : “the claims as a whole implement the alleged abstract idea into a practical application in a manner that imposes a meaningful limit on the abstract idea”, and the reasoning provided on page 8 paragraph 12 is: because again the “computing device” “establishing audio-video connections” … etc.
Respectfully when all the claim limitations and their outcome would not require any machine, in that case any claimed machine (e.g. the “computing device” here) becomes what is defined basically as unnecessary “additional elements”. Just merely mentioning various limitations are done by the “computing device”, does not make it an “improvement” and/or teaching “significantly more”. In order for the latter criteria to be satisfied as one example, it is not what is done by e.g. the “computing device” that would make it eligible, but rather how the limitations in reverse have had any impact on the “computing device” by e.g. resulting “faster search times”, and/or “smaller memory requirements” and/or perform a function by the computer element never done before by any computer element.
Regarding “Step 2B” a copy of the latest amendments is provided on page 8 the last paragraph and it is concluded that: “The Office fails to establish that the particular way of receiving and transmitting hand sign related video and audio is well understood”.
Respectfully the “Office” hand not examined the latest amendments which possess the amendments beginning with the “receiving” and the “transmitting” words. For that please visit the new office action.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-2, 4, 6-10, 12-20, and 22 stand rejected:
Claims 1, 12, and 17 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The claims recite how to “transmit[]” “audio stream comprising” “audio converted from the sign language hand signs” in an “audio-video connections” “communication session” from one “participant” of the “communication session” “to one or more second computing devices” according to “an emotional state of the [sending] participant” (e.g., spec. ¶ 0035 page 5 lines 1+: “For example, if the general mood of others” “is” “happy” “that contextual information” “may suggest” “happy voice model” “to annunciate signs made by Signer”).
Therefore, these claims limitations, as drafted, are processes that, under their broadest reasonable interpretations, cover performance of the limitations in the mind but for the recitation of a generic computer component (“computing device” (claim 1)). That is, other than reciting “by a computing device”, nothing in the claim limitations precludes their steps from practically being performed in the mind; i.e., a human who is not himself hearing impaired and also possesses vocal ability and is also trained in understanding “sign language”, could translate “sign language” associated with a hearing impaired person to other people either people proximate to him, or participants within a “video conference”. Furthermore, the “human” (the participant) with the knowledge of the “sign language” can also use inputs from any of the one or more other participants or members of the audience that is he translating the “sign language” for, to provide inputs to aid in the translations in multiple ways, e.g., impart specific feeling or emotions associated with the topic under discussion (using a specific voice model)).
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claims recite an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claims only recite one additional element, i.e., a “computing device” (in claim 1) to carry out all the claim limitations of “establishing” “receiving” “selecting” “converting” and “transmitting”. The said “computing device” though is recited at a high-level of generality (i.e., as a generic computing device performing all the claim limitations) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claims are thus directed to an abstract idea.
The claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using a “computing device” to perform the claim limitations amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic computer component cannot provide an inventive concept. The claims are thus not patent eligible.
Regarding claim 2, “selecting” “the voice model” “based on emotional state” requires observing a participant’s reactions to probe his emotions which is an operation based on observation which is one of the 101 ineligible operations.
Regarding claim 4, the method basically depends on analyzing “video” (i.e., image) which amounts to an observation process which is identified as patent ineligible.
Regarding claim 6, determining a participant “emotional state” based on “meeting mood level” (e.g. speed) is an action doable by the human making signs who can simply adjust the speed of presenting the hand signs to be slower and/or faster.
Regarding claims 7 and 16, as each word in a sign language is associated with a specific hand sign and/or hand signs, therefore the human translator will use automatically a sequence of hand signs without needing any machines.
Regarding claims 8 and 18, they depend on analysis of “participant” “facial expression[]” which is an observation process and thus patent ineligible.
Regarding claims 9 and 17, the human translator may present for each specific hand gesture associated with each word voice with different tones in voice to e.g. convey different accents.
Regarding claim 10, “video embellishment” detection is an analysis based on “identif[ying]” “hand signs” which is an observation process and thus patent ineligible.
Regarding claim 13, the human translator could also convey an ambient even while translating hand signs.
Regarding claims 14 and 19, the human translator (the participant) could also convey his environmental conditions while translating his or her hand signs.
Regarding claim 15, “selecting” a “voice model based on emotional states of one or more other participants” would merely require the human to direct his sign language to voice translation towards any one other than someone initially he was engaged with.
Regarding claim 20, the human translator could elect to translate the hand signs of the signer into textual format. This does not require anything beyond a pen and a paper.
Regarding claim 22, the human (presenter) could make his presentations based on a collective response (mood information) he gets from the participants, e.g. speeding up and/or slowing down in presentations depending on how he perceives the audience is grasping it. The inclusion of “deep neural network” is at such generality that it does not set any meaningful limitations on the claim and therefore serve as insignificant extra solution activity, as it is also a well known technique; see also spec. ¶ 0065 last S: “A multi-input, multi-layer deep neural network may be used to combine the various emotional mood indicators” which neither teaches any specific known neural network and/or any neural network generated based on the instant application requirements.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-2, 4, 6-10 stand rejected:
Claim 1 recites the limitation "the plurality of different voice models" in the 3rd limitation. There is insufficient antecedent basis for this limitation in the claim.
Regarding claims 2, 4, 6-10 as they depend on claim 1 and as they do not obviate the problem noted in their parent claim, they are thus rejected under similar rationale.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-2, 4, 6-10 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Weber (US 2019/0171716).
Regarding claim 1, Weber does teach a method (Title, Abstract)
comprising:
establishing by a first computing device, one or more audio-video connections with one or more second computing devices in a communication session (¶ 0028 S1: “Server” (a first computing device) “supports” “audio” “video” (establishes audio-video connections) “chat”, which according to ¶ 0025 S1: “server” (the first computing device) “renders” (is in communication with) “multimedia content including but not limited to the composite audio/video stream to each of the participants to the video conference” (in a communication session) “via one or more user interfaces” (to one or more secondary computing devices));
receiving, via an audio-video connection to the one or more audio-video connections, sign language hand signs from a participant of the communication session (step “604”: “capture” (receive) “by a video camera of the first endpoint, a video stream” (in an audio video connection) “of the first message expressed in sign language” (sign language hand signs) “by the user” (from a participant); Abstract lines 1+: “A user interface” “is presented for a hearing-impaired user” (a participant’s presentation) “to make selections that influence how that user participates in a video conference” (to be received in the communication session) “and how other participants” (comprising a plurality of other participants) “in that video conference interact with the hearing-impaired user”);
selecting a voice model from [a] an emotional state of the participant (¶ 0026 lines 6+: “For example, hearing-impaired users may select” (selecting) “qualities such as gender, accent, speed, pronunciation, prosody” (based on an emotional state of the participant from a plurality of voice models e.g. one associated with selection based on “gender” another based on “accent” …etc.) “so as to emulate to a desirable video conference persona”);
converting, using the selected voice model, the sign language hand signs to audio; and transmitting, to the one or more second computing devices and via the one or more audio-video connections, an audio stream comprising the audio converted from the signa language hand signs (Abstract last sentence: “associated with the sign-language-to-speech” (converting sign language hand signs to audio) “translations, with speech signals being produced” “so as to include certain effects” (based on e.g., “prosody” (the emotional state)) “when played out” (the generated “speech” (the converted audio)) “at the endpoints of the other participants” (to be transmitted to the video or second computing devices of each participant) “or emulate” “video conference persona” (e.g., using a model based on “prosody” (emotional state)) “from the standpoint of the hearing-impaired user”).
Regarding claim 2, Weber does teach the method of claim 1, wherein the selecting the voice model is further based on emotional states of one or more other participants in the communication session (¶ 0028 last sentence: “Alternatively, or in addition, endpoints” (each participant (e.g. comprising of one or more other participants in the “video conference” (communication session)) possesses) “that incorporate other sensors, such as data gloves, accelerometers, or other hand, finger, facial expression” (a model based on emotional state) “and/or body movement tracking sensors may be used, and the data collected thereby provided as part of the feed to server 108 for use in translating” (for converting observed associated signs) “the sign language to text and/or to speech” (to voice)).
Regarding claim 4, Weber does teach the method of claim 1, further comprising identifying the emotional state of the participant based on video from a camera associated with the participant (¶ 0029 sentence 1: “In some embodiments, a hearing-impaired user can join a video conference with a first device” (using a first “camera” (¶ 0028)) “that is capable of supporting video and a second device” (and using a second “camera”) “that is better suited for providing a data feed from one or more sensors used to capture sign language gestures, facial expressions” (to capture the emotional state second video images in addition to the “first device” (“camera”); ¶ 0028 lines 14-15: “users” “may use conventional web cameras” (using camera) “and the like to capture sign language gestures” (to capture sign language gestures for the conversion to “speech” (to aid voice models)); ¶ 0036 lines 13+: “the user may need to be coached to perform the signing within the field of view of one or more cameras” (using plurality of cameras for the signer)).
Regarding claim 6, Weber does teach the method of claim 1, wherein the emotional state of the participant is based on a meeting mood level (¶ 0036 lines 15+: “Also, the user may be asked to sign” (image which is used in probing “facial expression” (emotional state)) “at slower than usual speeds” (i.e., is based on “speed” (a mood level) that may demand slower speech) “so as to ensure all of his/her words are captured and understood” (so other participants can better comprehend (i.e., for the plurality of participants)); ¶ 0026 sentence 2: “For example, hearing-impaired users may select” (select voice model) “qualities such as gender, accent, speed” (based on the meeting mood level) “pronunciation, prosody, and/or other linguistic characteristics, so as to emulate a desirable video conference persona from the standpoint of the hearing-impaired user”).
Regarding claim 7, Weber does teach the method of claim 1, further comprising selecting different voice models to annunciate different words in a sequence of sign language hand signs (¶ 0033 column 2 lines 2+: “sign language capture system 502 includes one or more processors configured to remove noise from captured image sequences” (using sequence of) “such as a transient image of a human hand” (hand signs associated with the “sign language” (sign language)) “in the act of signing” (for the purpose of “sign-language-to-speech” (audio annunciations)) and in so doing according to ¶ 0026 sentence 2+: “ For example, hearing-impaired users may select” (user selects) “qualities such as gender, accent, speed, pronunciation, prosody, and/or other linguistic characteristics” (different voice models) “To that end, the video conferencing system may be provisioned with a stored user profile” (e.g., the “profile” uses one model for “accent” and another model for “gender” for playback in the video conference of the said “sequence” (sequence of signs)) “for a hearing-impaired user that includes previous selections of the user with respect to such items”).
Regarding claim 8, Weber does teach the method of claim 1, wherein the emotional state of the participant is based on one or more of:
a facial expression of the participant (¶ 0033 lines 1+: “Sign language capture system 502 is configured to capture hand and finger gestures and shapes, and, optionally, facial expressions” (selecting facial expressions (emotional state)) “and/or body movements, made by a hearing-impaired user” (of the participant) “and provide a video and/or data feed representing those captures” “to video conferencing system 100” (for conversion of hand signs to voice, because: ¶0025 lines 11+: “In general, such translations” “consider[] and recogniz[e] finger and/or hand shapes” “and/or” “facial expressions” (facial expressions used for) “for translating sign language to” “speech” (to aid in conversion of hand signs to voice (i.e. functions as voice model));
a facial expression of another participant of a plurality of participants of the communication session;
a reaction of another user who is not among the plurality of participants; and
information indicating an excitement level of a portion of a video stream being viewed by the plurality of participants.
Regarding claim 9, Weber does teach the method of claim 1, wherein the plurality of different voice models comprise different audio annunciations of a first hand sign (¶ 0034 last sentence: “In such instances, the hearing-impaired user may determine the nature of the audio signals to be so played, for example, selecting playback qualities such as gender, accent” (using associated different voice models) “etc. that will influence how the hearing-impaired user's “voice”” (for a specific sign hand) “will be played out” (play back for audio annunciations) “at the other participants' respective endpoints” (e.g., the “hearing-impaired user’s “voice”” (e.g. a specific hand sign associated with a word can be audibly annunciated in different “accent[s]” (example of audio annunciation) and/or different “gender[s]” (another audio annunciation example) and/or combinations thereof)).
Regarding claim 10, Weber does teach the method of claim 1, further comprising adding, based on the identified emotional state and the sign language hand signs, an audio or video embellishment to the communication session (¶ 0033 lines 1+: “Sign language capture system 502 is configured to capture hand and finger gestures and shapes, and, optionally, facial expressions and/or body movements” (examples of video embellishments) “made by a hearing-impaired user, and provide a video and/or data feed” (being added) “representing those captures to video conferencing system” (to the communication session)).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 12-18, 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over weber, and further in view of KYU et al. (KR 102045761B1).
Regarding claim 12, Weber does teach a method (Title, Abstract)
comprising:
storing a plurality of different voice models for annunciating a particular sign language sign (¶ 0026 sentence 2+: “For example, hearing-impaired users may select qualities such as gender, accent, speed, pronunciation, prosody” (a plurality of voice models) “and/or other linguistic characteristics, so as to emulate a desirable video conference persona from the standpoint of the hearing-impaired user” (for interpreting and annunciating a sign language) “In some cases, even idiosyncrasies such as the use of certain vocabulary can be selected through the user interface. To that end, the video conferencing system may be provisioned with a stored” (is being stored) “user profile for a hearing-impaired user that includes previous selections of the user with respect to such items”);
storing one or more voice model selection rules (¶ 0026 sentence 2+: “ For example, hearing-impaired users may select” (user selects) “qualities such as gender, accent, speed, pronunciation, prosody, and/or other linguistic characteristics” (using different voice models indicating different contexts) “To that end, the video conferencing system may be provisioned with a stored user profile” (according to a selection rule based on the signer’s contexts, e.g., the “profile” restricts a certain “accent” (one voice model (context)) in combination with a “gender” (a different voice model pertaining to that user context) …etc. for playback in the video conference) “for a hearing-impaired user that includes previous selections of the user with respect to such items”);
receiving, via an audio-video connection for a communication session, sign language hand signs from a participant of the communication session (step “604”: “capture” (receive) “by a video camera of the first endpoint, a video stream” (in an audio video connection) “of the first message expressed in sign language” (sign language hand signs) “by the user” (from a participant); Abstract lines 1+: “A user interface” “is presented for a hearing-impaired user” (a participant’s presentation) “to make selections that influence how that user participates in a video conference” (to be received in a communication session) “and how other participants” (comprising a plurality of other participants) “in that video conference interact with the hearing-impaired user”);
based on an emotional state and the one or more voice model selection rules, selecting a voice model of the plurality of different voice models (¶ 0026 lines 6+: “For example, hearing-impaired users may select” (selecting) “qualities such as gender, accent, speech, pronunciation, prosody” (based on an emotional state of the participant from a plurality of voice models e.g. one associated with selection based on “gender” another based on “accent” …etc.) “so as to emulate to a desirable video conference persona”);
converting, using the selected voice model, the sign language hand signs to audio; and transmitting, to the one or more second computing devices associated with the communication session, an audio stream comprising the audio converted from the sign language hand signs (Abstract last sentence: “associated with the sign-language-to-speech” (converting sign language hand signs to audio) “translations, with speech signals being produced” “so as to include certain effects” (based on e.g., “prosody” (the emotional state)) “when played out” (the generated “speech” (the converted audio)) “at the endpoints of the other participants” (to be transmitted to the video or second computing devices of each participant of the “video conference” (communication session)) “or emulate” “video conference persona” (e.g., using a model based on “prosody” (emotional state)) “from the standpoint of the hearing-impaired user”).
Weber does not specifically disclose:
Storing one or more voice model selection rules indicating associations between the plurality of different voice models and different emotional states.
Kyu et al. do teach:
Storing one or more voice model selection rules indicating associations between the plurality of different voice models and different emotional states (¶ 0047: “The” “voice synthesis model selection reference module 220 stores” (storing) “reference information” (selection rules) “for selecting a voice synthesis model suitable” (indicating an association between a voice model) “for each character and reference information for selecting a voice synthesis model suitable for each emotional state” (and an emotional state); ¶ 0048 last S: “each emotional state, such as ‘sadness’, ‘joy’, and ‘surprise’” (there are a plurality of different emotional states and each is associated with a voice model resulting in a plurality of different voice models)).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “voice synthesis model selection reference model” of Kyu et al. into the “video conference persona” of Weber would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Weber to expand its “video persona” “emulation[s]” to not only comprise of “persona” based on “prosody” but also of a plurality of other emotional states such as “sadness”, “joy”, “surprise” etc. as disclosed in Kyu et al. ¶ 0048.
Regarding claim 13, Weber does teach the method of claim 12, wherein the rules comprise rules for selecting the selected voice model based on content metadata indicating events occurring in a content item being viewed by the participant (¶ 0044: “For completeness, it is noted that in addition to transmitting the textual translation of the first message to the second endpoint, a video stream” (a video stream (content metadata or item)) “of the hearing-impaired user may also be transmitted to the second endpoint, so that participant(s)” (being viewed by the plurality of participants including the participant) “at the second endpoint may see the hearing-impaired user. This video stream may be identical to or different from the video stream that captures the first message expressed in sign language”; ¶ 0044 lines 21+: “In instances where the participant(s) at the second endpoint do not fully understand the communicated signs from the hearing-impaired user at the first endpoint, the participant(s) may rely upon the audio” (while receiving in voice the content associated with the video stream) “or textual translation of the signed message” (i.e., using a voice model based on viewing associated content in the video stream)).
Regarding claim 14, Weber does teach the method of claim 12, wherein the rules comprise rules for selecting the selected voice model based on environmental conditions of an environment of the participant (¶ 0033 2nd column lines 5+: “In certain embodiments, sign language capture system 502 is in communication with local display 406 via a local PC 404 (not shown) on a local area network (e.g., using wired (e.g., Ethernet) and/or wireless (e.g., Wi-Fi) connections). This allows the user to monitor the images” (i.e., in the environment of the signer he views his performance) “being transmitted so as to ensure they are being properly captured” (to decide whether or not it qualifies for sign to speech conversion (i.e., impacts selection of voice model)) “In certain embodiments, sign language capture system 502 is in communication with local display 406 via a video conferencing infrastructure such as video conferencing system 100”).
Regarding claim 15, Weber does teach the method of claim 12, wherein the rules comprise rules for selecting the selected voice model based on emotional states of one or more other participants in the video communication session (¶ 0026 lines 6+: “For example, hearing-impaired users” (each participant of the “video conference” (communication session)) “may select” (can select) “qualities such as gender, accent, speech, pronunciation, prosody” (based on his “prosody” (emotional state) a voice model) “so as to emulate to a desirable video conference persona”).
Regarding claim 16, Weber does teach the method of claim 12, further comprising using different voice models to annunciate different words in a sequence of signs, wherein the different voice models comprise different audio annunciations of a same word (¶ 0033 2nd column lines 2+: “sign language capture system 502 includes one or more processors configured to remove noise from captured image sequences” (using sequence of) “such as a transient image of a human hand” (hand signs) “in the act of signing” (for the purpose of “sign-language-to-speech” (audio annunciations)) and in so doing according to ¶ 0026 sentence 2+: “ For example, hearing-impaired users may select” (user selects) “qualities such as gender, accent, speed, pronunciation, prosody, and/or other linguistic characteristics” (different voice models) “To that end, the video conferencing system may be provisioned with a stored user profile” (e.g., the “profile” uses one model for “accent” and another model for “gender” for playback in the video conference of the said “sequence” (sequence of signs)) “for a hearing-impaired user that includes previous selections of the user with respect to such items”).
Regarding claim 17, Weber does teach a method (Title, Abstract)
comprising:
receiving, by a first computing device, video comprising a participant in a communication session (Abstract lines 1+: “A user interface” (a first computing device receives) “is presented for a hearing-impaired user” (a participant) “to make selections that influence how that user participates in a video conference” (in a communication session) “and how other participants” (and also comprising a plurality of other participants) “in that video conference interact with the hearing-impaired user”);
detecting, in the video, a sequence of sign language hand signs from the participant (¶ 0033 2nd column lines 2+: “sign language capture system 502 includes one or more processors configured to remove noise from captured image sequences” (using sequence of) “such as a transient image of a human hand” (sing language hand signs) “in the act of signing” (from the participant for the purpose of “sign-language-to-speech” (audio annunciations));
identifying emotional states of the participant while making different signs in the sequence (¶ 0026 lines 6+: “For example, hearing-impaired users may select” (identifying) “qualities such as gender, accent, speed, pronunciation, prosody” (“speed” and “prosody” (emotional states) of the participant) “so as to emulate to a desirable video conference persona” (while making different hand signs in sequence));
and
using different audio voice models to annunciate different signs in the sequence sign language of hand signs, wherein the different audio voice models each comprise a different audio annunciation for a same hand sign (¶ 0034 last sentence: “In such instances, the hearing-impaired user may determine the nature of the audio signals to be so played, for example, selecting playback qualities such as gender, accent” (using associated different voice models) “etc. that will influence how the hearing-impaired user's “voice”” (for a specific or same sign hand) “will be played out” (play back or audio annunciations) “at the other participants' respective endpoints” (e.g., the “hearing-impaired user’s “voice”” (e.g. a specific hand sign associated with a word can be audibly annunciated in different “accent[s]” (example of audio annunciation) and/or different “gender[s]” (another audio annunciation example) and/or combinations thereof)),
transmitting, to one or more second computing devices in the communication session and via one or more audio-video connections, an audio stream comprising audio annunciating the different signs (Abstract last sentence: “associated with the sign-language-to-speech” (annunciating different sign language hand signs to audio stream) “translations, with speech signals being produced” “so as to include certain effects” “when played out” “at the endpoints of the other participants” (to be transmitted to the video or second computing devices of each participant of the “video conference” (communication session or another audio-video connected device)) “or emulate” “video conference persona” “from the standpoint of the hearing-impaired user”).
Weber does not specifically disclose:
Wherein the different audio voice models are selected based on the corresponding identified emotional states.
Kyu et al. do teach:
Wherein the different audio voice models are selected based on the corresponding identified emotional states (¶ 0047: “The” “voice synthesis model selection reference module 220 stores” “reference information” “for selecting a voice synthesis model suitable” (selecting a voice model) “for each character and reference information for selecting a voice synthesis model suitable for each emotional state” (and a corresponding emotional state); ¶ 0048 last S: “each emotional state, such as ‘sadness’, ‘joy’, and ‘surprise’” (there are a plurality of different emotional states and each is associated with a voice model resulting in a plurality of different voice models)).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “voice synthesis model selection reference model” of Kyu et al. into the “video conference persona” of Weber would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Weber to expand its “video persona” “emulation[s]” to not only comprise of “persona” based on “prosody” but also of a plurality of other emotional states such as “sadness”, “joy”, “surprise” etc. as disclosed in Kyu et al. ¶ 0048.
Regarding claim 18, Weber does teach the method of claim 17, wherein the identifying emotional states is based on facial expressions of the participant while making the sequence of sign language hand signs (¶ 0033 lines 1+: “Sign language capture system 502 is configured to capture hand and finger gestures and shapes, and, optionally, facial expressions” (selecting facial expressions (emotional state)) “and/or body movements, made by a hearing-impaired user” (of the participant) “and provide a video and/or data feed representing those captures” “to video conferencing system 100” (for conversion of hand signs to voice, because: ¶0025 lines 11+: “In general, such translations” “consider[] and recogniz[e] finger and/or hand shapes” “and/or” “facial expressions” (facial expressions used for) “for translating sign language to” “speech” (to aid in conversion of hand signs to voice (i.e. functions as voice model)).
Regarding claim 20, Weber does teach the method of claim 17, further comprising translating the sequence of sign language hand signs to a textual transcript (¶ 0013: “FIG. 6 depicts a flow diagram of a process to capture, at a first endpoint, a video stream of a message expressed in sign language” (hand signs) “automatically translate” (converted) “the video stream of the message into a textual” (to text or textual transcript) “translation of the message,
and transmit the textual translation of the message to a second endpoint”).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Weber in view Kyu et al., and further in view of KARNATAKI et al. (US 2024/0282218).
Regarding claim 19, Weber in view of Kyu et al. do teach the method of claim 17, further comprising: and wherein the using the different audio voice models is further based on different environmental conditions associated with the video (Weber: ¶ 0036 last 7 lines: “the user interface for the sign language capture system 502” “include user-selectable setting for language of translation, signing convention being used (e.g., American Sign Language vs other conventions” (using different voice models depending on the environmental condition associated with the video and/or audio, e.g., selection of one language of translation versus another and/or one sign language convention versus another depending on the other participant’s language (audio) and/or sign (video) convention).
Weber in view of Kyu et al. do not specifically disclose the method of claim 17, further comprising:
monitoring environmental conditions, associated with the video, during the sequence of sign language hand signs.
KARNATAKI et al. do teach
monitoring environmental conditions, associated with the video, during the sequence of sign language hand signs (0014: “FIG. 1 illustrates working environment” (environmental conditions) “of a portable assistive device converting sign” (for sequence of hand signs associated with a video) “language to text and speech” “and vice versa in accordance with the present invention”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “environment” considerations of “converting sign” “to text and speech” of KARNATAKI et al. into the “sing-language to speech” of Weber in Weber in view of Kyu et al. would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable a signer to better prepare for the environment that he/or she is to present sign language such as using appropriate “gloves” as disclosed in KARNATAKI et al. ¶ 0051.
Claim(s) 22 is/are rejected under 35 U.S.C. 103 as being unpatentable over Weber, and further in view of Ray et al. (US 2024/0161373).
Regarding claim 22, Weber does not specifically disclose the method of claim 6, wherein a multi-input, multi-layer deep neural network is used for the combining the mood information for the plurality of participants.
Ray et al. do teach the method of claim 6, wherein a multi-input, multi-layer deep neural network is used for the combining the mood information for the plurality of participants (¶ 0019 lines 9-12 referring to Fig. 1 : “For example, the trained machine-learning model” (using a multi-layer multi-input deep neural network) “can determine a composite emotional state” (determines a combined mood information) “of the users 106a-b” (for the “user 106a” and “106b” (a plurality of participants)) “in the video conference” (in a video meeting) “call based on the input” where the “composite emotional state” (the combined mood information)).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “machine learning model” used for “video conference” of Ray et al. into the methods used for “video conference” of Weber would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Weber by virtue of the “machine learning” to determine “a composite emotional state of the users” in the “video conference” as disclosed in Ray et al. ¶ 0009 last sentence.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. KANAI (US 2020/0342896) Abstract lines 3+: “a storage part that stores” (storing) “a voice recognition model” (one or more voice model selection rules) “corresponding” (indicating association between a voice model) “to human emotion” (and a user emotional state). Here there exists only one “emotion” and one “model”.
As regards to allowability, claiming functions of “control codes” does help overcoming Alice 101 and teachings pertaining to “hand sign speed and size parameters” mapped to “different voice models” are determined novel. These were proposed as examiner’s amendments to the applicant representative, details of which are provided with an attached interview summary.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to FARZAD KAZEMINEZHAD whose telephone number is (571)270-5860. The examiner can normally be reached 10:30 am to 11:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached at (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Farzad Kazeminezhad/
Art Unit 2653
June 25th 2026.