Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Response to Amendment
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-4, 6, 8-11, 13, 15-18 and 21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Dharmarajan (2017/0277684) in view of Maizels et al (2024/0070251).
Consider claims 1, 8 and 15, Dharmarajan teaches a computer-implemented method, system and non-transitory machine-readable storage medium storing instructions that when executed by one or more processors, cause the one or more processors operations/comprising: extracting data from a video stream of a communication session, wherein the data includes a video frame and an audio segment, wherein the video frame includes a representation of a particular gesture (par. 0036-0037; “Also, each communication device 100, including the communication devices 100a and 100b, may incorporate and/or be coupled to one or more of a camera 110, an input device 120, an audio interface 170 and a display 180. The audio interface 170 may include and/or may enable coupling to one or more acoustic transducers (e.g., one or more microphones, one or more speakers, etc.) to capture and/or acoustically output sounds that may include speech”; par. 0040; 0070; “Alternatively or additionally, the processor 150 of the communication device 100b may operate one or more microphones incorporated into and/or coupled to the audio interface 170 to capture speech sounds uttered by the operator of the communication device 100b. The processor 150 may then be caused to operate the network interface 190 to transmit those captured speech sounds to the connection server 300 for conversion into the images of the avatar making sign language gestures as has been described”); executing a neural network using the data, wherein the neural network generates a predicted communication intended by the particular gesture; generating a communication response based on the predicted communication intended by the particular gesture (par. 0089-0090; “receive images captured of the operator who uses sign language making sign language gestures. As has been discussed, if the signing routine 210 is executed within a connection server 300 such that conversions to and from sign language are performed by the processor 350 thereof, then the captured images of the operator making sign language gestures may be received from the communication device of the operator who uses sign language via a network extending therebetween (e.g., a portion of the network 999)”; “the processor may employ training data earlier generated during a training operation to improve the accuracy of the interpretation of those sign language gestures as part of performing the conversion”); and facilitating a transmission of the communication response, the communication response being a response to the particular gesture (par. 0091; “the processor may transmit the text generated by this conversion to the other communication device of the other operator who does not use sign language via the network. As previously discussed, if the conversion from sign language to text is performed within the communication device of the operator who uses sign language, then the text may be relayed through the connection server, as well as through the network”).
Dharmarajan did not explicitly suggest wherein the audio segment includes non-verbal sound; executing a neural network using the data, wherein the neural network generates a predicted communication intended by the particular gesture and the audio segment; generating a communication response based on the predicted communication intended by the particular gesture and the audio segment. In the same field of endeavor, Maizels et al suggested such (par. 00175; “Consistent with the present disclosure, speech detection system 100 may be capable of detecting facial skin micromovements of user 102 and extract meaning from the detected movements, even without vocalization of speech or utterance of any other sounds by user 102. The extracted meaning may be an identification of user 102 wearing speech detection system 100, an identification of a subvocalization by a user, such as a word silently spoken by user 102, an identification of a word vocally spoken by user 102, an identification of a phoneme silently spoken by user 102, or an identification of a phoneme vocally spoken by user 102. Similarly, the extract meaning may include an identification of a heart rate of user 102, an identification of a breathing rate of user 102, and/or other characteristics associated with verbal or non-verbal communication by user 102. In one example, speech detection system 100 may generate output signals that include data associated with an identification information, a UI command, synthesized audio signal, a textual transcription, or any combination thereof”; par. 0282; “The term “non-verbal cues” refers to the various forms of communication that occur without the use of spoken words. Some examples of non-verbal cues may include facial expressions, body language, gestures, eye contact, tone of voice, postures, and other subtle signals that convey meaning in interpersonal interactions”).
Therefore, it would have been obvious to one of the ordinary skills in the art before the effective filing date to substitute the speech detection of Dharmarajan with Maizels et al speech detection and the results would have been predictable and resulted in detecting and identifying non-verbal speech thereby providing enhanced and improved communications for impaired user during a video call.
Consider claims 2, 9 and 16, Dharmarajan teaches wherein the communication session is between a user device and a terminal device (par. 0083; “a processor of a connection server of a communication system (e.g., the processor 350 of a connection servers 300 of the communication system 1000) may receive, via a network, a request from a communication device associated with an operator using sign language (e.g., the communication device 100a) to communicate with one of multiple operators of one of multiple communication devices”).
Consider claims 3, 10 and 17, Dharmarajan teaches wherein the particular gesture is a static position of a body part (par. 0026; “It should also be noted that the images captured of an operator making sign language gestures for subsequent conversion of sign language to text may include captured motion video (e.g., multiple images captured as motion video frames at a predetermined frame rate) of the operator forming the sign language gestures”).
Consider claim 4, 11 and 18, Dharmarajan teaches wherein the particular gesture is a motion involving one or more body parts (par. 0026; “It should also be noted that the images captured of an operator making sign language gestures for subsequent conversion of sign language to text may include captured motion video (e.g., multiple images captured as motion video frames at a predetermined frame rate) of the operator forming the sign language gestures”).
Consider claims 6, 13 and 21, Dharmarajan teaches wherein the neural network is an ensemble network comprising two or more neural networks configured to generate outputs of different types (par. 0030; “the first communication device or the connection server may operate in a training mode to learn idiosyncrasies of the first operator in making sign language gestures. In such a training mode, the first operator may make a number of sign language gestures of their choosing and/or may be guided through making a selection of sign language gestures within a field of view of a camera of the first communication device”; par. 0031; “Alternatively or additionally, the first communication device may transmit profile data to the connection server indicating that the first communication device includes at least a camera to capture sign language gestures and/or a display to present an avatar making sign language gestures. The connection server may employ such indications in profile data to trigger and/or enable conversions between sign language and text, and/or between sign language and speech as part of enabling the first operator to more easily converse with another operator who does not use sign language”).
Claims 7, 14 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Dharmarajan (2017/0277684) in view Maizels et al (2024/0070251) and further in view of Jandhyala et al (2022/0157083).
Consider claims 7, 14 and 20, Dharmarajan does not explicitly suggest wherein the neural network is configured to generate a boundary box over the particular gesture. In the same field of endeavor, Jandhyala et al teach a systems and method that identify a gesture based on event camera data and frame-based camera data. In some implementations, the frame-based camera data is used to identify a region of interest (e.g., a bounding box) for the event camera to analyze (par. 0003; 0059-0060; “a plurality of events 730 within the bounding box are detected by the event camera 422b. In some implementations, the events 730 detected by the event camera 422b include positive events (e.g., generated by the leading edge of the hand or fingers) and negative events (e.g., generated by the trailing edge of the hand of fingers)”). Therefore, it would have been obvious to a person of ordinary skills in the art before the effective filing date the invention was made to incorporate the teaching of Jandhyala et al into view of Dharmarajan and Maizels et al and the results would have been predictable and resulted in providing an improved process for identify gesture thereby allow the system to quickly, efficiently, and accurately identify and classify gestures.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-4, 6-11, 13-18 and 20-21 have been considered but are moot.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any response to this action should be mailed to:
Mail Stop ____(explanation, e.g., Amendment or After-final, etc.) Commissioner for Patents
P.O. Box 1450
Alexandria, VA 22313-1450
Facsimile responses should be faxed to:
(571) 273-8300
Hand-delivered responses should be brought to:
Customer Service Window
Randolph Building
401 Dulany Street
Alexandria, VA 22314
Any inquiry concerning this communication or earlier communications from the examiner should be directed to QUOC DUC TRAN whose telephone number is (571)272-7511. The examiner can normally be reached Monday-Friday 8:30am - 5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached on (571) 272-7503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Quoc D Tran/
Primary Examiner, Art Unit 2691
August 27, 2026