DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged of certified copies of papers submitted under 35 U.S.C. 119(a)-(d), which papers have been placed of record on file.
Information Disclosure Statement(s)
The Information disclosure statement (IDS) filed on February 13th, 2025 has been acknowledged and considered by the examiner.
Claim Objections
Claims 4, 12 and 15 are objected to because of the following informalities:
Claim 4, line 5, the reference of “MediaPipe” needs to be given specific definition, within the claim, for the reference to be definite and avoid 112(b) indefiniteness issue, unclear definition. Appropriate correction is required.
Claim 12, line 13, “execute the instruction” should be read as “execute the instruction mapped to the gesture” to follow proper antecedent basis reference back to the first instantiation of “an instruction mapped to the gesture” in lines 11-12, since there is instance of using the same term in line 3, “storing instructions”, therefore, the suggested amendment is to differentiate and give proper antecedent reference. Appropriate correction is required.
Claim 15, line 5, the reference of “MediaPipe” needs to be given specific definition for the reference to be definite and avoid 112(b) indefiniteness issue, unclear definition. Appropriate correction is required.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitation(s) that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that use the word “means” or “step” but are nonetheless not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph because the claim limitation(s) recite(s) sufficient structure, materials, or acts to entirely perform the recited function.
Claims 3-4, 6-7, 13-15 and 17-18, recite(s) limitation(s) that use words like “means” (or “step”) or similar terms with functional language and do invoke 35 U.S.C. 112(f):
Claim 3, recites the limitation, “the gesture recognition…to perform palm recognition…” [Lines 1-2].
Claim 4; recites the limitation, “a gesture recognition model for annotating hand landmark positions” [Line 5].
Claims 4 and 5; recites the limitation, “a gesture recognition model …obtain” [Line 5-7].
Claim 6, recites the limitation, “the pre-established gesture recognition model, to obtain a hand landmark position” [Lines 5-6].
Claim 7; recites the limitation, “an object detection model, to determine the hand shape change and the gesture motion…” [Lines 6-7].
Claim 13; recites the limitation, “a gesture recognition model to obtain the gesture recognition result image data…” [Lines 5-6].
Claim 14; recites the limitation, “the gesture recognition model…to perform palm recognition and palm landmark position recognition” [Lines 1-2].
Claim 15; recites the limitation, “a gesture recognition model for annotating hand landmark positions” [Lines 5-6].
Claim 17; recites the limitation, “the pre-established gesture recognition model, to obtain a hand landmark position” [Lines 4-5].
Claim 18; recites the limitation, “an object detection model, to determine the hand shape change...” [Lines 5-6].
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
After a careful analysis, as disclosed above, and a careful review of the specification the following limitations in claims 3-4, 6-7, 13-15 and 17-18;
(i) “gesture recognition model”, there is insufficient support in the disclosure, filed on December 06th, 2024, for the recited gesture recognition model to perform the recited functions, the closest disclosure can be found in Par. [0063] of the disclosure, wherein the model is based on MediaPipe and trained using a training set, which is not sufficient to provide specific support for a structure, material and/or act for the recited gesture recognition model, thus have no sufficient structure or material and/or act).
(ii) “object detection model”, there is no sufficient support for the recited object detection model to perform the recited functions, the closest disclosure can be found in the specification, filed on December 06th, 2024, in Par. [0041], wherein the object detection model is used in conjunction with the gesture recognition model which is nominal, there is no sufficient support of a structure, material and/or act for the recited object detection model to perform the recited functions, thus have no sufficient structure or material and/or act.).
(iii) “pre-established gesture recognition model” there is no sufficient support for the recited pre-established gesture recognition model to perform the recited functions, the closest disclosure can be found in the specification, filed on December 06th, 2024, in Par. [0036], wherein the pre-established gesture recognition model is the gesture recognition model after being pre-established to perform obtaining of the gesture recognition result image data, which is nominal, there is no sufficient support of a structure, material and/or act for the recited object detection model to perform the recited functions, thus have no sufficient structure or material and/or act.).
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 3-4, 7, 13-15 and 17-18 along with their associated/dependent claims are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
Claim 3-4, 6-7, 13-15 and 17-18’s limitations:
Claim 3, recites the limitation, “the gesture recognition…to perform palm recognition…” [Lines 1-2].
Claim 4; recites the limitation, “a gesture recognition model for annotating hand landmark positions” [Line 5].
Claims 4 and 5; recites the limitation, “a gesture recognition model …obtain” [Line 5-7].
Claim 6, recites the limitation, “the pre-established gesture recognition model, to obtain a hand landmark position” [Lines 5-6].
Claim 7; recites the limitation, “an object detection model, to determine the hand shape change and the gesture motion…” [Lines 6-7].
Claim 13; recites the limitation, “a gesture recognition model to obtain the gesture recognition result image data…” [Lines 5-6].
Claim 14; recites the limitation, “the gesture recognition model…to perform palm recognition and palm landmark position recognition” [Lines 1-2].
Claim 15; recites the limitation, “a gesture recognition model for annotating hand landmark positions” [Lines 5-6].
Claim 17; recites the limitation, “the pre-established gesture recognition model, to obtain a hand landmark position” [Lines 4-5].
Claim 18; recites the limitation, “an object detection model, to determine the hand shape change...” [Lines 5-6].
Claim 3-4, 6-7, 13-15 and 17-18, each respectively invokes 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function. The specification is devoid of adequate structure to perform the claimed functions. The specification does not provide sufficient details such that one of the ordinary skill in the art would understand which structure performed(s) the claimed function.
Therefore, the claim is indefinite and is rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph.
Applicant may:
(a) Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph;
(b) Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
(c) Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)).
If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either:
(a) Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
(b) Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181.
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claim 3-4, 6-7, 13-15 and 17-18 along with their associated/dependent claims are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for pre-AIA the inventor(s), at the time the application was filed, had possession of the claimed invention. As described above, the disclosure does not provide adequate structure to perform the claimed function in the recited limitation.
Claim 3-4, 6-7, 13-15 and 17-18’s limitations:
Claim 3, recites the limitation, “the gesture recognition…to perform palm recognition…” [Lines 1-2].
Claim 4; recites the limitation, “a gesture recognition model for annotating hand landmark positions” [Line 5].
Claims 4 and 5; recites the limitation, “a gesture recognition model …obtain” [Line 5-7].
Claim 6, recites the limitation, “the pre-established gesture recognition model, to obtain a hand landmark position” [Lines 5-6].
Claim 7; recites the limitation, “an object detection model, to determine the hand shape change and the gesture motion…” [Lines 6-7].
Claim 13; recites the limitation, “a gesture recognition model to obtain the gesture recognition result image data…” [Lines 5-6].
Claim 14; recites the limitation, “the gesture recognition model…to perform palm recognition and palm landmark position recognition” [Lines 1-2].
Claim 15; recites the limitation, “a gesture recognition model for annotating hand landmark positions” [Lines 5-6].
Claim 17; recites the limitation, “the pre-established gesture recognition model, to obtain a hand landmark position” [Lines 4-5].
Claim 18; recites the limitation, “an object detection model, to determine the hand shape change...” [Lines 5-6].
The specification does not demonstrate that applicant has made an invention that achieves the claimed function because the invention is not described with sufficient detail such that one of ordinary skill in the art can reasonably conclude that the inventor had possession of the claimed invention.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Xiaofeng Tong et. al. (“US 2013/0278504 A1” hereinafter as “Tong”) in view of Bernd Ette et. al. (“US 2020/0143154 A1” hereinafter as “Ette”).
Regarding claim 1, Tong teaches an interaction processing method (Abstract discloses “…recognize a dynamic hand gesture and provide corresponding user interface command”), comprising: receiving a dynamic image of a gesture move of a user (Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture in response to depth images of gesture”; Par. [0024] discloses “track hand region through a sequence of depth images (e.g., a sequence of video frames captured by camera”, sequence of images in a video is analogous to the recited dynamic image); performing gesture recognition on the dynamic image (Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture” which is analogous to the recited gesture recognition [gesture recognition engine]) to obtain gesture recognition result image data of the dynamic image (Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture in response to depth images of gesture” wherein the dynamic hand gesture is analogous to the recited gesture recognition result image data; moreover, Par. [0020] discloses “…dynamic hand gesture including one or more hand postures combined with various hand motion trajectories”); to determine a hand shape change (Par. [0030] discloses “gesture recognition module may also include gesture identification module generally configured to identify one or more shape features of the hand”; furthermore, Par. [0031] discloses “…shape features may include…difference between left and right portions of a bounding box, and/or the difference between top and bottom portions of a bounding box” indicating bounding box detection of the hand shape indicating difference in the detection, being analogous to a hand shape change as claimed, wherein Par. [0041] discloses “…a move gesture including a hand changing posture”) and a gesture motion trajectory of the user (Par. [0030] discloses “…identify one or more shape features of the hand in binary image generated by hand tracking module and to use those shape features and the motion trajectory information”); determining, based on the hand shape change and the gesture motion trajectory, a gesture (Par. [0030] discloses “use those hand features and the motion trajectory information provided by hand tracking module to identify dynamic hand gestures”) corresponding to the hand shape change and the gesture motion trajectory (Par. [0030] discloses “use those hand features and the motion trajectory information provided by hand tracking module to identify dynamic hand gestures”; moreover, Par. [0039] discloses “…a dynamic hand gesture based on the identified hand posture in combination with hand motion trajectory information”) and an instruction mapped to the gesture (Par. [0042] discloses “…that gesture identification provide gesture commands in response to the recognition of dynamic hand gestures” indicating an instruction [gesture commands]; furthermore, Par. [0041] discloses “…compare the dynamic hang gesture’s hand posture and motion feature to a database of hand postures and motion features of predefined dynamic hand gestures. For example, FIG. 5 depicts various predefined dynamic hand gestures and corresponding UI commands” a database of hand gesture and corresponding commands indicating mapping of corresponding information); and executing the instruction (Par. [0044] discloses “…used by gesture command module to generate UI command that may be recognized by application to effect the corresponding UI action”).
However, Tong does not explicitly teach performing object detection based on the gesture recognition result image data, to determine a hand shape change and a gesture motion trajectory.
Ette teaches performing object detection based on the gesture recognition result image data (Pars. [0125-0126] disclose “a frame object has been detected within a frame of the image data…the outline of the frame object is parameterized to produce an interval profile”; furthermore, Pars. [0129-0131] disclose “the interval profile can be used to identify a maximum corresponding to an extended finger…a gesture can already be detected on the basis of the interval profile…the interval profile can be compared with one or more reference profiles…one or more profiles associated with the gesture detected in the present case, to be updated and stored”), to determine a hand shape change and a gesture motion trajectory (Par. [0035] discloses “changes on the basis of the shape of the frame objects of successive frames can be checked for plausibility”; Par. [0120] discloses “detect a gesture on the basis of an interval profile and/or a trajectory”; furthermore, Par. [0133] discloses “the curve shape can be analyzed and, for example, the interval between the maxima associated with the extended fingers can be evaluated” wherein, the detection of the hand and its gesture can be used to further identify fingers and their maxima as part of an object detection to determine interval profile of a complete hand gesture as taught by Ette).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong of having an interaction processing method, comprising: receiving a dynamic image of a gesture move of a user; performing gesture recognition on the dynamic image to obtain gesture recognition result image data of the dynamic image; to determine a hand shape change and a gesture motion trajectory of the user; determining, based on the hand shape change and the gesture motion trajectory, a gesture corresponding to the hand shape change and the gesture motion trajectory and an instruction mapped to the gesture; and executing the instruction, with the teachings of Ette of having wherein performing object detection based on the gesture recognition result image data, to determine a hand shape change and a gesture motion trajectory.
Wherein having Tong’s method wherein having performing object detection based on the gesture recognition result image data, to determine a hand shape change and a gesture motion trajectory.
The motivation behind the modification would have been to perform gesture detection on improving user interfaces and user interaction with hand movement interactive devices, and further to perform reliable detection of a capture gesture and reduce computation power. Since both Tong and Ette’s systems perform hand gesture detection. Wherein Tong’s system improves gesture detection and hand tracking for executing commands that enable improvement on user interfaces and user interaction with hand movement interactive devices (see Tong’s Par. [0001]), and Ette’s system improves performing gesture detection for a reliable detection of a capture gesture and reduce computation power (see Ette’ Par. [0014]).
Regarding claim 12, Tong teaches an interaction processing apparatus (Abstract discloses “…recognize a dynamic hand gesture and provide corresponding user interface command”), comprising: a processor; and a memory storing instructions executable by the processor, wherein the processor is configured to (Par. [0013] discloses “instructions stored on a machine-readable medium, which may be read and executed by one or more processors”): receive a dynamic image of a gesture move of a user (Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture in response to depth images of gesture” ; Par. [0024] discloses “track hand region through a sequence of depth images (e.g., a sequence of video frames captured by camera”, sequence of images in a video is analogous to the recited dynamic image); perform gesture recognition on the dynamic image (Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture” which is analogous to the recited gesture recognition [gesture recognition engine]) to obtain gesture recognition result image data of the dynamic image (Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture in response to depth images of gesture” wherein the dynamic hand gesture is analogous to the recited gesture recognition result image data; moreover, Par. [0020] discloses “…dynamic hand gesture including one or more hand postures combined with various hand motion trajectories”); to determine a hand shape change (Par. [0030] discloses “gesture recognition module may also include gesture identification module generally configured to identify one or more shape features of the hand”; furthermore, Par. [0031] discloses “…shape features may include…difference between left and right portions of a bounding box, and/or the difference between top and bottom portions of a bounding box” indicating bounding box detection of the hand shape indicating difference in the detection, being analogous to a hand shape change as claimed, wherein Par. [0041] discloses “…a move gesture including a hand changing posture”) and a gesture motion trajectory of the user (Par. [0030] discloses “…identify one or more shape features of the hand in binary image generated by hand tracking module and to use those shape features and the motion trajectory information”); determine, based on the hand shape change and the gesture motion trajectory, a gesture (Par. [0030] discloses “use those hand features and the motion trajectory information provided by hand tracking module to identify dynamic hand gestures”) corresponding to the hand shape change and the gesture motion trajectory (Par. [0030] discloses “use those hand features and the motion trajectory information provided by hand tracking module to identify dynamic hand gestures”; moreover, Par. [0039] discloses “…a dynamic hand gesture based on the identified hand posture in combination with hand motion trajectory information”) and an instruction mapped to the gesture (Par. [0042] discloses “…that gesture identification provide gesture commands in response to the recognition of dynamic hand gestures” indicating an instruction [gesture commands]; furthermore, Par. [0041] discloses “…compare the dynamic hang gesture’s hand posture and motion feature to a database of hand postures and motion features of predefined dynamic hand gestures. For example, FIG. 5 depicts various predefined dynamic hand gestures and corresponding UI commands” a database of hand gesture and corresponding commands indicating mapping of corresponding information); and execute the instruction (Par. [0044] discloses “…used by gesture command module to generate UI command that may be recognized by application to effect the corresponding UI action”).
However, Tong does not explicitly teach to perform object detection based on the gesture recognition result image data, to determine a hand shape change and a gesture motion trajectory.
Ette teaches to perform object detection based on the gesture recognition result image data (Pars. [0125-0126] disclose “a frame object has been detected within a frame of the image data…the outline of the frame object is parameterized to produce an interval profile”; furthermore, Pars. [0129-0131] disclose “the interval profile can be used to identify a maximum corresponding to an extended finger…a gesture can already be detected on the basis of the interval profile…the interval profile can be compared with one or more reference profiles…one or more profiles associated with the gesture detected in the present case, to be updated and stored”), to determine a hand shape change and a gesture motion trajectory (Par. [0035] discloses “changes on the basis of the shape of the frame objects of successive frames can be checked for plausibility”; Par. [0120] discloses “detect a gesture on the basis of an interval profile and/or a trajectory”; furthermore, Par. [0133] discloses “the curve shape can be analyzed and, for example, the interval between the maxima associated with the extended fingers can be evaluated” wherein, the detection of the hand and its gesture can be used to further identify fingers and their maxima as part of an object detection to determine interval profile of a complete hand gesture as taught by Ette).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong of having an interaction apparatus to perform instructions stored in a memory to be executed by a processor comprises: receive a dynamic image of a gesture move of a user; perform gesture recognition on the dynamic image to obtain gesture recognition result image data of the dynamic image; to determine a hand shape change and a gesture motion trajectory of the user; determine, based on the hand shape change and the gesture motion trajectory, a gesture corresponding to the hand shape change and the gesture motion trajectory and an instruction mapped to the gesture; and execute the instruction, with the teachings of Ette of having wherein performing object detection based on the gesture recognition result image data, to determine a hand shape change and a gesture motion trajectory.
Wherein Tong’s apparatus wherein having performing object detection based on the gesture recognition result image data, to determine a hand shape change and a gesture motion trajectory.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform reliable detection of a capture gesture and reduce computation power. Since both Tong and Ette’s systems perform hand gesture detection. Wherein Tong’s system improves gesture detection and hand tracking for executing commands that enable improvement on user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Ette’s system improves performing gesture detection for a reliable detection of a capture gesture and reduce computation power (see Ette’ Par. [0014]).
Claims 2 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Xiaofeng Tong et. al. (“US 2013/0278504 A1” hereinafter as “Tong”) in view of Bernd Ette et. al. (“US 2020/0143154 A1” hereinafter as “Ette”) and Hung-Quoc Duc Lai et. al. (“US 2019/0327124 A1” hereinafter as “Lai”).
Regarding claim 2, Tong in view of Ette teaches the interaction processing method according to claim 1, Tong teaches further comprising: performing preprocessing on the dynamic image to obtain a processed dynamic image (Par. [0114] discloses “a preprocessing can be performed for the reference profiles and/or the reference trajectories…preprocessing can be effected such that suitable features can already be extracted in advance and provided for the profile comparison”; Par. [0106] discloses “features extracted from the frames can be detected…trajectory features are extracted” hence, the frames are dynamic image, and their extracted features are being preprocessed on); and the performing the gesture recognition on the dynamic image (Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture” which is analogous to the recited gesture recognition [gesture recognition engine]) to obtain the gesture recognition result image data of the dynamic image (Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture in response to depth images of gesture” wherein the dynamic hand gesture is analogous to the recited gesture recognition result image data; moreover, Par. [0020] discloses “…dynamic hand gesture including one or more hand postures combined with various hand motion trajectories”) comprises: inputting the processed dynamic image (Par. [0067] discloses “data required for the profile comparison…the extreme values, may already have been preprocessed and stored for the reference interval profiles”) into a gesture recognition model (Par. [0070] discloses “profile comparison is performed on the basis of a machine learning method, for example, on the basis of a neural network”, the preprocessed image data being used for a profile comparison which is a neural network [analogous to the recited gesture recognition model]) to obtain the gesture recognition result image data (Par. [0118] discloses “after a gesture has been determined on the basis of the captured image data, a post-processing is performed in a further operation” wherein, the determined gesture is analogous to the recited gesture recognition result).
However, Tong in view of Ette does not explicitly teach performing image transformation preprocessing on the dynamic image to obtain a processed dynamic image.
Lai teaches performing image transformation preprocessing on the dynamic image to obtain a processed dynamic image (Par. [0461] discloses “image preprocessing may be applied, including…transformation…feature extraction”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette of having an interaction processing method, comprising performing preprocessing on the dynamic image to obtain a processed dynamic image; and the performing the gesture recognition on the dynamic image to obtain the gesture recognition result image data of the dynamic image comprises: inputting the processed dynamic image into a gesture recognition model to obtain the gesture recognition result image, with the teachings of Lai of having wherein performing image transformation preprocessing on the dynamic image to obtain a processed dynamic image
Wherein, Tong’s method wherein having performing image transformation preprocessing on the dynamic image to obtain a processed dynamic image.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform correcting error on the image and offset compensation on an image through transformation on the image for better effective feature extraction. Since both Tong and Lai’s systems perform hand gesture detection and tracking. Wherein Tong’s system improves gesture detection and hand tracking for executing commands that enable improvement on user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Lai’s system improves on feature extraction by correcting error on the image and performing offset compensation on an image through transformation for better effective feature extraction (see Lai’s Par. [0534]).
Regarding claim 13, Tong in view of Ette teaches the interaction processing apparatus according to claim 12, Tong teaches wherein the processor is further configured to: perform preprocessing on the dynamic image to obtain a processed dynamic image (Par. [0114] discloses “a preprocessing can be performed for the reference profiles and/or the reference trajectories…preprocessing can be effected such that suitable features can already be extracted in advance and provided for the profile comparison”; Par. [0106] discloses “features extracted from the frames can be detected…trajectory features are extracted” hence, the frames are dynamic image, and their extracted features are being preprocessed on); and input the processed dynamic image (Par. [0067] discloses “data required for the profile comparison…the extreme values, may already have been preprocessed and stored for the reference interval profiles”) into a gesture recognition model (Par. [0070] discloses “profile comparison is performed on the basis of a machine learning method, for example, on the basis of a neural network”, the preprocessed image data being used for a profile comparison which is a neural network [analogous to the recited gesture recognition model]) to obtain the gesture recognition result image data (Par. [0118] discloses “after a gesture has been determined on the basis of the captured image data, a post-processing is performed in a further operation” wherein, the determined gesture is analogous to the recited gesture recognition result).
However, Tong in view of Ette does not explicitly teach to perform image transformation preprocessing on the dynamic image to obtain a processed dynamic image.
Lai teaches to perform image transformation preprocessing on the dynamic image to obtain a processed dynamic image (Par. [0461] discloses “image preprocessing may be applied, including…transformation…feature extraction”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette of having an interaction processing apparatus to execute instructions, stored in a memory to be executed by a processor, comprises: perform preprocessing on the dynamic image to obtain a processed dynamic image; and input the processed dynamic image into a gesture recognition model to obtain the gesture recognition result image, with the teachings of Lai of having wherein the processor is configured to perform image transformation preprocessing on the dynamic image to obtain a processed dynamic image
Wherein Tong’s apparatus, wherein the processor is configured to perform image transformation preprocessing on the dynamic image to obtain a processed dynamic image.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform correcting error on the image and offset compensation on an image through transformation on the preprocessing of the image for better effective feature extraction. Since both Tong and Lai’s systems perform hand gesture detection. Wherein Tong’s system improves gesture detection and hand tracking for executing commands that enable improvement on user interfaces and user interaction with hand gesture interactive devices (Tong’s Par. [0001]), and Lai’s system improves on feature extraction by correcting error on the image and performing offset compensation on an image through transformation for better effective feature extraction (Lai’s Par. [0534]).
Claims 3 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Xiaofeng Tong et. al. (“US 2013/0278504 A1” hereinafter as “Tong”) in view of Bernd Ette et. al. (“US 2020/0143154 A1” hereinafter as “Ette”) further in view of Hung-Quoc Duc Lai et. al. (“US 2019/0327124 A1” hereinafter as “Lai”) and Ayan Sinha et. al. (“US 2023/0044664 A1” hereinafter as “Sinha”).
Regarding claim 3, Tong in view of Ette and Lai teaches the interaction processing method according to claim 2.
However, Tong in view of Ette and Lai does not explicitly teach wherein the gesture recognition model is pre-established to perform palm recognition and palm landmark position recognition on an input image to obtain a gesture recognition result.
Sinha teaches wherein the gesture recognition model is pre-established (Par. [0023] discloses “engine uses the activation feature data from the neural networks to search…”hand pose parameter”…refer to any data that describe a portion of the pose of a hand, such as the relative location and angle of a given joint in the hand…the orientation and shape of the palm”, the neural network here is analogous to the recited gesture recognition model; Par. [0006] discloses “discriminative features learned from deep convolutional neural nets to recognize the hand pose…track the hand model for robust gesture intent recognition”) to perform palm recognition (Par. [0023] discloses “engine uses the activation feature data from the neural networks to search…”hand pose parameter”…refer to any data that describe a portion of the pose of a hand, such as the relative location and angle of a given joint in the hand…the orientation and shape of the palm” indicating a palm recognition) and palm landmark position recognition on an input image to obtain a gesture recognition result (Par. [0005] discloses “processing hand gestures in a three-dimensional space is to identify the pose of the hand…the pose is affected by the rotation of the wrist, the shape of the palm, and the positions of the fingers…” furthermore, Par. [0018] discloses “wrist refers to the joint at the base of the hand formed form the carpal bones that connects to hand to the forearm and to the metacarpals bones that for the palm”, wherein, the palm including joint information and orientation and shape of the palm, is analogous to palm’s landmark).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette and Lai of having an interaction processing method, comprising performing preprocessing on the dynamic image to obtain a processed dynamic image; and the performing the gesture recognition on the dynamic image to obtain the gesture recognition result image data of the dynamic image, with the teachings of Sinha of having wherein the gesture recognition model is pre-established to perform palm recognition and palm landmark position recognition on an input image to obtain a gesture recognition result.
Wherein Tong’s method having wherein the gesture recognition model is pre-established to perform palm recognition and palm landmark position recognition on an input image to obtain a gesture recognition result.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to take into consideration of different hand features during feature extraction for enabling greater range of gesture movements of the gesture recognition enable improvement on operation of hand gesture recognition engines. Since both Tong and Sinha’s systems perform hand gesture detection. Wherein Tong’s system improves gesture detection and hand tracking for executing commands that enable improvement on user interfaces and user interaction with hand gesture interactive devices (Tong’s Par. [0001]), and Sinha’s system improves hand gesture recognition by taking into consideration of different hand features during feature extraction to enable greater range of gesture movements for gesture recognition and improves on operation of hand gesture recognition engines (Sinha’s Par. [0004] and Par. [0006]).
Regarding claim 14, Tong in view of Ette and Lai teaches the interaction processing apparatus according to claim 13.
However, Tong in view of Ette and Lai does not explicitly teach wherein the gesture recognition model is pre-established to perform palm recognition and palm landmark position recognition on an input image to obtain a gesture recognition result.
Sinha teaches wherein the gesture recognition model is pre-established (Par. [0023] discloses “engine uses the activation feature data from the neural networks to search…”hand pose parameter”…refer to any data that describe a portion of the pose of a hand, such as the relative location and angle of a given joint in the hand…the orientation and shape of the palm”, the neural network here is analogous to the recited gesture recognition model; Par. [0006] discloses “discriminative features learned from deep convolutional neural nets to recognize the hand pose…track the hand model for robust gesture intent recognition”) to perform palm recognition (Par. [0023] discloses “engine uses the activation feature data from the neural networks to search…”hand pose parameter”…refer to any data that describe a portion of the pose of a hand, such as the relative location and angle of a given joint in the hand…the orientation and shape of the palm” indicating a palm recognition) and palm landmark position recognition on an input image to obtain a gesture recognition result (Par. [0005] discloses “processing hand gestures in a three-dimensional space is to identify the pose of the hand…the pose is affected by the rotation of the wrist, the shape of the palm, and the positions of the fingers…” furthermore, Par. [0018] discloses “wrist refers to the joint at the base of the hand formed form the carpal bones that connects to hand to the forearm and to the metacarpals bones that for the palm”, wherein, the palm including joint information and orientation and shape of the palm, is analogous to palm’s landmark).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette and Lai of having an interaction processing apparatus to execute instructions comprises: perform preprocessing on the dynamic image to obtain a processed dynamic image; and the performing the gesture recognition on the dynamic image to obtain the gesture recognition result image data of the dynamic image, with the teachings of Sinha of having wherein the gesture recognition model is pre-established to perform palm recognition and palm landmark position recognition on an input image to obtain a gesture recognition result.
Wherein Tong’s apparatus having the gesture recognition model is pre-established to perform palm recognition and palm landmark position recognition on an input image to obtain a gesture recognition result.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to take into consideration of different hand features during feature extraction to enable greater range of gesture movements for gesture recognition and improve on operation of hand gesture recognition engines. Since both Tong and Sinha’s systems perform hand gesture detection. Wherein Tong’s system improves gesture detection and hand tracking for executing commands that enable improvement on user interfaces and user interaction with hand gesture interactive devices (Tong’s Par. [0001]), and Sinha’s system improves on hand gesture recognition by taking into consideration of different hand features during feature extraction to enable greater range of gesture movements for gesture recognition and improve operation of hand gesture recognition engines (Sinha’s Par. [0004] and Par. [0006]).
Claims 4-5 and 15-16 are rejected under 35 U.S.C. 103 as being unpatentable over Xiaofeng Tong et. al. (“US 2013/0278504 A1” hereinafter as “Tong”) in view of Bernd Ette et. al. (“US 2020/0143154 A1” hereinafter as “Ette”) further in view of Hung-Quoc Duc Lai et. al. (“US 2019/0327124 A1” hereinafter as “Lai”) and Ayan Sinha et. al. (“US 2023/0044664 A1” hereinafter as “Sinha”) and Sylvana Alpert et. al. (“US 12,530,087 B1” hereinafter as “Alpert”).
Regarding claim 4, Tong in view of Ette further in view of Lai and Sinha teaches the interaction processing method according to claim 3.
However, Tong in view of Ette further in view of Lai does not explicitly teach wherein pre-establishing the gesture recognition model comprises: obtaining a plurality of gesture images, and performing hand area annotation to form a training set; and training the built gesture recognition model using the training set, to obtain the trained gesture recognition model.
Sinha teaches wherein pre-establishing the gesture recognition model comprises (Par. [0006] discloses “discriminative features learned from deep convolutional neural nets to recognize the hand pose…for robust gesture intent recognition…training the network”, wherein the learning/training indicates a pre-establishing of the model): obtaining a plurality of gesture images (Par. [0006] discloses “training the network on a large database of synthetic hands rendered from different camera line of sight”), and performing hand area annotation to form a training set (Par. [0050] discloses “synthetic training data including a plurality of frames …training data depth map that each correspond to one pose of the hand with known hand pose parameters”; Par. [0023] discloses “hand pose parameters…refer to any data that describe a portion of the pose of a hand…or any other description of the shape and orientation of the hand” indicating hand area annotation [data description]); and training the built gesture recognition model using the training set (Par. [0020] discloses training the neural network using the obtained training data), to obtain the trained gesture recognition model (Par. [0024] discloses “a process to train a neural network and recommendation engine to perform the hand pose identification process”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette and Lai of having an interaction processing method, comprising performing preprocessing on the dynamic image to obtain a processed dynamic image; and the performing the gesture recognition on the dynamic image to obtain the gesture recognition result image data of the dynamic image, with the teachings of Sinha of having wherein pre-establishing the gesture recognition model comprises: obtaining a plurality of gesture images, and performing hand area annotation to form a training set; and training the built gesture recognition model using the training set, to obtain the trained gesture recognition mode
Wherein Tong’s method having pre-establishing the gesture recognition model comprises: obtaining a plurality of gesture images, and performing hand area annotation to form a training set; and training the built gesture recognition model using the training set, to obtain the trained gesture recognition mode
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to take into consideration of different hand features during feature extraction to enable greater range of gesture movements for gesture recognition and improve operation of hand gesture recognition engines. Since both Tong and Sinha’s systems perform hand gesture detection. Wherein Tong’s system improves gesture detection and hand tracking for executing commands that enable improvement on user interfaces and user interaction with hand gesture interactive devices (Tong’s Par. [0001]), and Sinha’s system improves hand gesture recognition by taking into consideration of different hand features during feature extraction to enable greater range of gesture movements for gesture recognition and improve operation of hand gesture recognition engines (Sinha’s Par. [0004] and Par. [0006]).
However, Tong in view of Ette further in view of Lai and Sinha does not explicitly teach performing hand landmark annotation to form a training set; building, based on MediaPipe, a gesture recognition model for annotating hand landmark positions in an image.
Alpert teaches performing hand landmark annotation to form a training set (Col. 7, lines 1-11, discloses “the detected hand regions are input in hand landmark model, which is configured to determine landmarks in the hand regions…hand landmark model outputs landmark locations and confidence score for each landmark”; furthermore, Col.8, lines 23-39, discloses “location of each hand region in the first frame (e.g., x, y coordinates defining a hand region..” indicating annotation of the hand landmark [coordinates defining the region]); building, based on MediaPipe, a gesture recognition model (Col. 8, lines 1-17, discloses “an example gesture recognition algorithm is the gesture algorithm included in the MediaPipe framework”) for annotating hand landmark positions in an image (Col. 8, lines 1-17, discloses “gesture recognizer, which detects a hand or finer gesture and places gesture recognition data (e.g., a gesture descriptor)” indicating hand positions, furthermore, Col. 7, lines 1-11, discloses “the detected hand regions are input in hand landmark model, which is configured to determine landmarks in the hand regions”; Col.8, lines 23-39, discloses “location of each hand region in the first frame (e.g., x, y coordinates defining a hand region..”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette and Lai and Sinha of having a pre-establishing the gesture recognition model comprises: obtaining a plurality of gesture images, and performing hand area annotation to form a training set; and training the built gesture recognition model using the training set, to obtain the trained gesture recognition mode, with the teachings of Alpert of wherein having performing hand landmark annotation to form a training set; building, based on MediaPipe, a gesture recognition model for annotating hand landmark positions in an image.
Wherein, Tong’s method of wherein performing and hand landmark annotation to form a training set; building, based on MediaPipe, a gesture recognition model for annotating hand landmark positions in an image.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform detect multiple instances of hands at the same and correctly detect multiple instances of the hands when they are close together. Since both Tong and Alpert’s systems perform hand gesture detection. Wherein Tong’s system improves gesture detection and hand tracking for executing commands that enable improvement on user interfaces and user interaction with hand gesture interactive devices (Tong’s Par. [0001]), and Alpert’s system improves on detecting multiple instances of hands at the same time and correctly detecting multiple instances of the hands when they are close together which enables efficient hand tracking (Alpert’s Col. 3, lines 1-5).
Regarding claim 5, Tong in view of Ette further in view of Lai and Sinha and Alpert teaches the interaction processing method according to claim 4.
However, Tong in view of Ette further in view of Lai and Alpert does not explicitly teach wherein pre-establishing the gesture recognition model further comprises: performing image augmentation on the plurality of gesture images to obtain an expanded training set; and the training the built gesture recognition model using the training set, to obtain the trained gesture recognition model comprises: training the built gesture recognition model using the expanded training set, to obtain the trained gesture recognition model.
Sinha teaches wherein pre-establishing the gesture recognition model further comprises (Par. [0006] discloses “discriminative features learned from deep convolutional neural nets to recognize the hand pose…for robust gesture intent recognition…training the network”, wherein the learning/training indicates a pre-establishing of the model): performing image augmentation on the plurality of gesture images to obtain an expanded training set (Par. [0006] discloses “training the network on a large database of synthetic hands rendered from different camera line of sight”; Par. [0050] discloses “generation of synthetic training data including a plurality of frames of training depth map data that correspond to a synthetically generated hand in a wide range of predetermined poses”, synthetic generation of training data from a hand model and generate a population of synthetic hand poses is analogous to image augmentation as claimed); and the training the built gesture recognition model using the training set (Par. [0006] discloses “training the network on a large database of synthetic hands rendered from different camera line of sight”), to obtain the trained gesture recognition model comprises (Par. [0024] discloses “a process to train a neural network and recommendation engine to perform the hand pose identification process”): training the built gesture recognition model using the expanded training set (Par. [0006] discloses “training the network on a large database of synthetic hands rendered from different camera line of sight”), to obtain the trained gesture recognition model (Par. [0006] discloses “training the network on a large database of synthetic hands rendered from different camera line of sight”; Par. [0050] discloses “generation of synthetic training data including a plurality of frames of training depth map data that correspond to a synthetically generated hand in a wide range of predetermined poses”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette and Lai and Alpert of having an interaction processing method, comprising performing preprocessing on the dynamic image to obtain a processed dynamic image; and the performing the gesture recognition on the dynamic image to obtain the gesture recognition result image data of the dynamic image, with the teachings of Sinha of having wherein pre-establishing the gesture recognition model further comprises: performing image augmentation on the plurality of gesture images to obtain an expanded training set; and the training the built gesture recognition model using the training set, to obtain the trained gesture recognition model comprises: training the built gesture recognition model using the expanded training set, to obtain the trained gesture recognition model.
Wherein Tong’s method having pre-establishing the gesture recognition model further comprises: performing image augmentation on the plurality of gesture images to obtain an expanded training set; and the training the built gesture recognition model using the training set, to obtain the trained gesture recognition model comprises: training the built gesture recognition model using the expanded training set, to obtain the trained gesture recognition model.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to take into consideration different hand features to enable greater range of gesture movements for gesture recognition and improve operation of hand gesture recognition engines. Since both Tong and Sinha’s systems perform hand gesture detection and tracking. Wherein Tong’s system improves gesture detection and hand tracking for executing commands that enable improvement on user interfaces and user interaction with hand gesture interactive devices (Tong’s Par. [0001]), and Sinha’s system improves hand gesture recognition by taking into consideration of different hand features during feature extraction to enable greater range of gesture movements for gesture recognition and improve operation of hand gesture recognition engines (Sinha’s Par. [0004] and Par. [0006]).
Regarding claim 15, Tong in view of Ette further in view of Lai and Sinha teaches the interaction processing apparatus according to claim 14.
However, Tong in view of Ette further in view of Lai does not explicitly teach wherein the processor is further configured to: obtain a plurality of gesture images, and perform hand area annotation to form a training set; and train the built gesture recognition model using the training set, to obtain the trained gesture recognition model.
Sinha teaches wherein the processor is further configured to: (Par. [0006] discloses “discriminative features learned from deep convolutional neural nets to recognize the hand pose…for robust gesture intent recognition…training the network”, wherein the learning/training indicates a pre-establishing of the model): obtain a plurality of gesture images (Par. [0006] discloses “training the network on a large database of synthetic hands rendered from different camera line of sight”), and perform hand area annotation to form a training set (Par. [0050] discloses “synthetic training data including a plurality of frames …training data depth map that each correspond to one pose of the hand with known hand pose parameters”; Par. [0023] discloses “hand pose parameters…refer to any data that describe a portion of the pose of a hand…or any other description of the shape and orientation of the hand” indicating hand area annotation [data description]); and train the built gesture recognition model using the training set (Par. [0020] discloses training the neural network using the obtained training data), to obtain the trained gesture recognition model (Par. [0024] discloses “a process to train a neural network and recommendation engine to perform the hand pose identification process”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette and Lai of having an interaction processing apparatus to execute instructions comprises: perform preprocessing on the dynamic image to obtain a processed dynamic image, with the teachings of Sinha of having wherein the processor is further configured to: obtain a plurality of gesture images, and perform hand area annotation to form a training set; and train the built gesture recognition model using the training set, to obtain the trained gesture recognition mode
Wherein, Tong’s apparatus having the processor is further configured to: obtain a plurality of gesture images, and perform hand area annotation to form a training set; and train the built gesture recognition model using the training set, to obtain the trained gesture recognition mode
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to take into consideration of different hand features during feature extraction to enable greater range of gesture movements for gesture recognition and improve operation of hand gesture recognition engines. Since both Tong and Sinha’s systems perform hand gesture detection. Wherein Tong’s system improves gesture detection and hand tracking for executing commands that enable improvement on user interfaces and user interaction with hand gesture interactive devices (Tong’s Par. [0001]), and Sinha’s system improves hand gesture recognition by taking into consideration of different hand features during feature extraction to enable greater range of gesture movements for gesture recognition and improve operation of hand gesture recognition engines (Sinha’s Par. [0004] and Par. [0006]).
However, Tong in view of Ette further in view of Lai and Sinha does not explicitly teach to perform hand landmark annotation to form a training set; build, based on MediaPipe, a gesture recognition model for annotating hand landmark positions in an image.
Alpert teaches to perform hand landmark annotation to form a training set (Col. 7, lines 1-11, discloses “the detected hand regions are input in hand landmark model, which is configured to determine landmarks in the hand regions…hand landmark model outputs landmark locations and confidence score for each landmark”; furthermore, Col.8, lines 23-39, discloses “location of each hand region in the first frame (e.g., x, y coordinates defining a hand region..” indicating annotation of the hand landmark [coordinates defining the region]); build, based on MediaPipe, a gesture recognition model (Col. 8, lines 1-17, discloses “an example gesture recognition algorithm is the gesture algorithm included in the MediaPipe framework”) for annotating hand landmark positions in an image (Col. 8, lines 1-17, discloses “gesture recognizer, which detects a hand or finer gesture and places gesture recognition data (e.g., a gesture descriptor)” indicating hand positions, furthermore, Col. 7, lines 1-11, discloses “the detected hand regions are input in hand landmark model, which is configured to determine landmarks in the hand regions”; Col.8, lines 23-39, discloses “location of each hand region in the first frame (e.g., x, y coordinates defining a hand region..”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette and Lai and Sinha of having an interaction processing apparatus to perform instructions comprises: obtain a plurality of gesture images, and perform hand area annotation to form a training set; and train the built gesture recognition model using the training set, to obtain the trained gesture recognition mode, with the teachings of Alpert of having performing hand landmark annotation to form a training set; building, based on MediaPipe, a gesture recognition model for annotating hand landmark positions in an image.
Wherein, Tong’s apparatus of having wherein performing hand landmark annotation to form a training set; building, based on MediaPipe, a gesture recognition model for annotating hand landmark positions in an image.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform detect multiple instances of hands at the same and correctly detect multiple instances of the hands when they are close together. Since both Tong and Alpert’s systems perform hand gesture detection. Wherein Tong’s system improves gesture detection and hand tracking for executing commands that enable improvement on user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Alpert’s system improves on detecting multiple instances of hands at the same time and correctly detect multiple instances of the hands when they are close together enable efficient hand tracking (see Alpert’s Col. 3, lines 1-5).
Regarding claim 16, Tong in view of Ette further in view of Lai and Sinha and Alpert teaches the interaction processing apparatus according to claim 15.
However, Tong in view of Ette further in view of Lai and Alpert does not explicitly teach wherein the processor is further configured to: perform image augmentation on the plurality of gesture images to obtain an expanded training set; train the built gesture recognition model using the expanded training set, to obtain the trained gesture recognition model.
Sinha teaches wherein the processor is further configured to: (Par. [0006] discloses “discriminative features learned from deep convolutional neural nets to recognize the hand pose…for robust gesture intent recognition…training the network”, wherein the learning/training indicates a pre-establishing of the model): perform image augmentation on the plurality of gesture images to obtain an expanded training set (Par. [0006] discloses “training the network on a large database of synthetic hands rendered from different camera line of sight”; Par. [0050] discloses “generation of synthetic training data including a plurality of frames of training depth map data that correspond to a synthetically generated hand in a wide range of predetermined poses”, synthetic generation of training data from a hand model and generate a population of synthetic hand poses is analogous to image augmentation as claimed); and train the built gesture recognition model using the expanded training set (Par. [0006] discloses “training the network on a large database of synthetic hands rendered from different camera line of sight”), to obtain the trained gesture recognition model (Par. [0006] discloses “training the network on a large database of synthetic hands rendered from different camera line of sight”; Par. [0050] discloses “generation of synthetic training data including a plurality of frames of training depth map data that correspond to a synthetically generated hand in a wide range of predetermined poses”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette further in view of Lai and Alpert of having an interaction processing apparatus to execute instructions comprises: perform preprocessing on the dynamic image to obtain a processed dynamic image, with the teachings of Sinha of having wherein the processor is configured to: perform image augmentation on the plurality of gesture images to obtain an expanded training set; and train the built gesture recognition model using the expanded training set, to obtain the trained gesture recognition model.
Wherein Tong’s apparatus having the processor is configured to: perform image augmentation on the plurality of gesture images to obtain an expanded training set; and train the built gesture recognition model using the expanded training set, to obtain the trained gesture recognition model.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to take into consideration of different hand features during feature extraction to enable greater range of gesture movements for gesture recognition and improve operation of hand gesture recognition engines. Since both Tong and Sinha’s systems perform hand gesture detection. Wherein Tong’s system improves gesture detection and hand tracking for executing commands that enable improvement on user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Sinha’s system improves hand gesture recognition by taking into consideration of different hand features during feature extraction to enable greater range of gesture movements for gesture recognition and improve operation of hand gesture recognition engines (see Sinha’s Par. [0004] and Par. [0006]).
Claims 6 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Xiaofeng Tong et. al. (“US 2013/0278504 A1” hereinafter as “Tong”) in view of Bernd Ette et. al. (“US 2020/0143154 A1” hereinafter as “Ette”) further in view of Hung-Quoc Duc Lai et. al. (“US 2019/0327124 A1” hereinafter as “Lai”) and Ayan Sinha et. al. (“US 2023/0044664 A1” hereinafter as “Sinha”) and Sylvana Alpert et. al. (“US 12,530,087 B1” hereinafter as “Alpert”) and Jonathan Marsden et. al. (“US 2023/0214458 A1” hereinafter as “Marsden”).
Regarding claim 6, Tong in view of Ette further in view of Lai and Sinha and Alpert teaches the interaction processing method according to claim 4, Tong teaches wherein the inputting the processed dynamic image (Par. [0067] discloses “data required for the profile comparison…the extreme values, may already have been preprocessed and stored for the reference interval profiles”) into the gesture recognition model (Par. [0070] discloses “profile comparison is performed on the basis of a machine learning method, for example, on the basis of a neural network”, the preprocessed image data being used for a profile comparison which is a neural network [analogous to the recited gesture recognition model]) to obtain the gesture recognition result image data (Par. [0118] discloses “after a gesture has been determined on the basis of the captured image data, a post-processing is performed in a further operation” wherein, the determined gesture is analogous to the recited gesture recognition result).
However, Tong in view of Ette further in view of Lai and Sinha does not explicitly teach inputting the plurality of frames of images into the pre-established gesture recognition model, to obtain a hand landmark position annotation result image of each frame of image.
Alpert teaches inputting the plurality of frames of images into the pre-established gesture recognition model (Col. 8, lines 1-17, discloses “an example gesture recognition algorithm is the gesture algorithm included in the MediaPipe framework”), to obtain a hand landmark position annotation result image of each frame of image (Col. 8, lines 1-17, discloses “gesture recognizer, which detects a hand or finer gesture and places gesture recognition data (e.g., a gesture descriptor)” indicating hand positions, furthermore, Col. 7, lines 1-11, discloses “the detected hand regions are input in hand landmark model, which is configured to determine landmarks in the hand regions”; Col.8, lines 23-39, discloses “location of each hand region in the first frame (e.g., x, y coordinates defining a hand region..”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette and Lai and Sinha of having an interaction processing method, wherein the inputting the processed dynamic image into the gesture recognition model to obtain the gesture recognition result image data, with the teachings of Alpert of having wherein inputting the plurality of frames of images into the pre-established gesture recognition model, to obtain a hand landmark position annotation result image of each frame of image
Wherein, Tong’s method of wherein inputting the plurality of frames of images into the pre-established gesture recognition model, to obtain a hand landmark position annotation result image of each frame of image
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform detect multiple instances of hands at the same and correctly detect multiple instances of the hands when they are close together. Since both Tong and Alpert’s systems perform hand gesture detection. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Alpert’s system detecting multiple instances of hands at the same and correctly detect multiple instances of the hands when they are close together (see Alpert’s Col. 3, lines 1-5).
However, Tong in view of Ette further in view of Lai and Sinha and Alpert does not explicitly teach splitting the processed dynamic image into a plurality of frames of images in temporal order; and arranging hand landmark position annotation result images of the plurality of frame images in temporal order, to obtain the gesture recognition result image data of the processed dynamic image.
Marsden teaches splitting the processed dynamic image into a plurality of frames of images in temporal order (Par. [0329] discloses “the temporal generalist neural networks and the temporal specialist neural networks are trained using a series of simulated hand position images that are temporally linked as gesture sequences representing real world hand gestures…temporally linked in previous frames to generated a current set of hand position parameters” indicating a plurality of frames in temporal order from a simulation [processed dynamic image] as input into the neural networks); and arranging hand landmark position annotation result images of the plurality of frame images in temporal order (Par. [0329] discloses “combination of a current simulated hand position image and a series of prior estimated hand position parameters temporally linked in previous frames to generate a current set of hand position parameters” indicating the hand position parameters [annotation/labeling according to Par. 0333] being arranged in temporal order [temporally linked]; moreover, Par. [0333] discloses “hand positions for training of neural network systems…specify a range of hand positions and position sequences, a range of hand anatomies, including palm size, fattiness, stubbiness, and skin tone and a range of backgrounds” indicating hand landmark positions), to obtain the gesture recognition result image data of the processed dynamic image (Pars. [0328-0329] disclose “set of base features identified as implementations such as convolutional neural network…gesture recognition…gesture sequences representing real world hand gestures” indicating gesture recognition result to be obtained from the neural networks).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette and Lai and Sinha and Alpert of having an interaction processing method, wherein the inputting the processed dynamic image into the gesture recognition model to obtain the gesture recognition result image data, with the teachings of Marsden of having wherein splitting the processed dynamic image into a plurality of frames of images in temporal order; and arranging hand landmark position annotation result images of the plurality of frame images in temporal order, to obtain the gesture recognition result image data of the processed dynamic image.
Wherein, Tong’s method of having wherein splitting the processed dynamic image into a plurality of frames of images in temporal order; and arranging hand landmark position annotation result images of the plurality of frame images in temporal order, to obtain the gesture recognition result image data of the processed dynamic image.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to improve gesture recognition using hand detection and hand pose estimation simulation for training of neural networks. Since both Tong and Marsden’s systems perform hand gesture detection. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with devices (see Tong’s Par. [0001]), and Marsden’s system improves gesture recognition using hand detection and hand pose estimation simulation for training of neural networks (see Marsden’s Par. [0078]).
Regarding claim 17, Tong in view of Ette further in view of Lai and Sinha and Alpert teaches the interaction processing apparatus according to claim 15.
However, Tong in view of Ette further in view of Lai and Sinha does not explicitly teach wherein the processor is further configured to: input the plurality of frames of images into the pre-established gesture recognition model, to obtain a hand landmark position annotation result image of each frame of image.
Alpert teaches the processor is further configured to: input the plurality of frames of images into the pre-established gesture recognition model (Col. 8, lines 1-17, discloses “an example gesture recognition algorithm is the gesture algorithm included in the MediaPipe framework”), to obtain a hand landmark position annotation result image of each frame of image (Col. 8, lines 1-17, discloses “gesture recognizer, which detects a hand or finer gesture and places gesture recognition data (e.g., a gesture descriptor)” indicating hand positions, furthermore, Col. 7, lines 1-11, discloses “the detected hand regions are input in hand landmark model, which is configured to determine landmarks in the hand regions”; Col.8, lines 23-39, discloses “location of each hand region in the first frame (e.g., x, y coordinates defining a hand region..”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette and Lai and Sinha of having an interaction processing apparatus, with the teachings of Alpert of having wherein the processor is configured to perform: input the plurality of frames of images into the pre-established gesture recognition model, to obtain a hand landmark position annotation result image of each frame of image
Wherein, Tong’s apparatus of having the processor is configured to perform: input the plurality of frames of images into the pre-established gesture recognition model, to obtain a hand landmark position annotation result image of each frame of image
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform detect multiple instances of hands at the same and correctly detect multiple instances of the hands when they are close together. Since both Tong and Alpert’s systems perform hand gesture detection. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Alpert’s system improves on detecting multiple instances of hands at the same and correctly detecting multiple instances of the hands when they are close together (see Alpert’s Col. 3, lines 1-5).
However, Tong in view of Ette further in view of Lai and Sinha and Alpert does not explicitly teach to split the processed dynamic image into a plurality of frames of images in temporal order; and arrange hand landmark position annotation result images of the plurality of frame images in temporal order, to obtain the gesture recognition result image data of the processed dynamic image.
Marsden teaches to split the processed dynamic image into a plurality of frames of images in temporal order (Par. [0329] discloses “the temporal generalist neural networks and the temporal specialist neural networks are trained using a series of simulated hand position images that are temporally linked as gesture sequences representing real world hand gestures…temporally linked in previous frames to generated a current set of hand position parameters” indicating a plurality of frames in temporal order from a simulation [processed dynamic image] as input into the neural networks); and arrange hand landmark position annotation result images of the plurality of frame images in temporal order (Par. [0329] discloses “combination of a current simulated hand position image and a series of prior estimated hand position parameters temporally linked in previous frames to generate a current set of hand position parameters” indicating the hand position parameters [annotation/labeling according to Par. 0333] being arranged in temporal order [temporally linked]; moreover, Par. [0333] discloses “hand positions for training of neural network systems…specify a range of hand positions and position sequences, a range of hand anatomies, including palm size, fattiness, stubbiness, and skin tone and a range of backgrounds” indicating hand landmark positions), to obtain the gesture recognition result image data of the processed dynamic image (Pars. [0328-0329] disclose “set of base features identified as implementations such as convolutional neural network…gesture recognition…gesture sequences representing real world hand gestures” indicating gesture recognition result to be obtained from the neural networks).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette and Lai and Sinha and Alpert of having an interaction processing apparatus, with the teachings of Marsden of having wherein splitting the processed dynamic image into a plurality of frames of images in temporal order; and arranging hand landmark position annotation result images of the plurality of frame images in temporal order, to obtain the gesture recognition result image data of the processed dynamic image.
Wherein, Tong’s apparatus of having wherein splitting the processed dynamic image into a plurality of frames of images in temporal order; and arranging hand landmark position annotation result images of the plurality of frame images in temporal order, to obtain the gesture recognition result image data of the processed dynamic image.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with devices, and further to improve gesture recognition using hand detection and hand pose estimation simulation for training of neural networks. Since both Tong and Marsden’s systems perform hand gesture detection. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, (see Tong’s Par. [0001]), and Marsden’s system improves gesture recognition using hand detection and hand pose estimation simulation for training of neural networks (see Marsden’s Par. [0078]).
Claims 7 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Xiaofeng Tong et. al. (“US 2013/0278504 A1” hereinafter as “Tong”) in view of Bernd Ette et. al. (“US 2020/0143154 A1” hereinafter as “Ette”) and Yang Zhou et. al. (“US 2022/0277595 A1” hereinafter as “Zhou”).
Regarding claim 7, Tong in view of Ette discloses the interaction processing method according to claim 1.
However, Tong does not explicitly teach performing object detection based on the gesture recognition result image data, to determine a hand shape change and a gesture motion trajectory comprises: inputting the sampled gesture recognition result image data into an object detection model, to determine the hand shape change and the gesture motion trajectory of the user.
Ette teaches performing object detection based on the gesture recognition result image data (Pars. [0125-0126] disclose “a frame object has been detected within a frame of the image data…the outline of the frame object is parameterized to produce an interval profile”; furthermore, Pars. [0129-0131] disclose “the interval profile can be used to identify a maximum corresponding to an extended finger…a gesture can already be detected on the basis of the interval profile…the interval profile can be compared with one or more reference profiles…one or more profiles associated with the gesture detected in the present case, to be updated and stored”), to determine a hand shape change and a gesture motion trajectory (Par. [0035] discloses “changes on the basis of the shape of the frame objects of successive frames can be checked for plausibility”; Par. [0120] discloses “detect a gesture on the basis of an interval profile and/or a trajectory”; furthermore, Par. [0133] discloses “the curve shape can be analyzed and, for example, the interval between the maxima associated with the extended fingers can be evaluated”) comprises: inputting the gesture recognition result image data into an object detection model (Par. [0070] discloses “the profile comparison is performed on the basis of a machine learning method…the classification and detection of a gesture on the basis of the trajectory, the interval profile and the profile comparison” indicating a neural network for gesture detection [object detection]), to determine the hand shape change and the gesture motion trajectory of the user (Par. [0035] discloses “changes on the basis of the shape of the frame objects of successive frames can be checked for plausibility”; Par. [0120] discloses “detect a gesture on the basis of an interval profile and/or a trajectory”; furthermore, Par. [0133] discloses “the curve shape can be analyzed and, for example, the interval between the maxima associated with the extended fingers can be evaluated” wherein, the detection of the hand and its gesture can be used to further identify fingers and their maxima as part of an object detection to determine interval profile of a complete hand gesture as taught by Ette).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong of having an interaction processing method, comprising: receiving a dynamic image of a gesture move of a user; performing gesture recognition on the dynamic image to obtain gesture recognition result image data of the dynamic image; to determine a hand shape change and a gesture motion trajectory of the user; determining, based on the hand shape change and the gesture motion trajectory, a gesture corresponding to the hand shape change and the gesture motion trajectory and an instruction mapped to the gesture; and executing the instruction, with the teachings of Ette of having wherein performing object detection based on the gesture recognition result image data, to determine a hand shape change and a gesture motion trajectory comprises: inputting the sampled gesture recognition result image data into an object detection model, to determine the hand shape change and the gesture motion trajectory of the user.
Wherein Tong’s method having performing object detection based on the gesture recognition result image data, to determine a hand shape change and a gesture motion trajectory comprises inputting the sampled gesture recognition result image data into an object detection model, to determine the hand shape change and the gesture motion trajectory of the user.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform reliable detection of a capture gesture and reduce computation power. Since both Tong and Ette’s systems perform hand gesture detection. Wherein Tong’s system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Ette’s system improves performing on a reliable detection of a capture gesture and reduce computation power (see Ette’ Par. [0014]).
However, Tong in view of Ette does not explicitly teach the performing the object detection further comprising: performing sampling on the gesture recognition result image data through frame extraction, to obtain sampled gesture recognition result image data; inputting the sampled gesture recognition result image data into an object detection model.
Zhou teaches the performing the object detection further comprising: performing sampling on the gesture recognition result image data through frame extraction (Par. [0063] discloses “determining the target bounding box based on the initial bounding boxes, the hand gesture detection apparatus…perform down-sampling processing on the initial bounding boxes to obtain spare bounding boxes”, wherein, the bounding boxes are gesture recognition result image data, down-sampling processing is analogous to sampling as recited, wherein the down-sampling is performed on the bounding boxes being feature extraction from the frame hence, is analogous to “through frame extraction”), to obtain sampled gesture recognition result image data (Par. [0064] discloses “the hand gesture detection apparatus selects the target bounding box from the spare bounding boxes obtained by the down-sampling processing”); inputting the sampled gesture recognition result image data into an object detection model (Par. [0064] discloses “the hand gesture detection apparatus selects the target bounding box from the spare bounding boxes obtained by the down-sampling processing”; furthermore, Par. [0075] discloses “the hand to be detected by using the gesture estimation model to obtain the gesture detection result of the hand to be detected”, wherein the gesture estimation model for detection of the hand is analogous to the object detection model as claimed).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Tong in view of Ette of having an interaction processing method comprises performing object detection based on the gesture recognition result image data, to determine a hand shape change and a gesture motion trajectory comprises: inputting the sampled gesture recognition result image data into an object detection model, to determine the hand shape change and the gesture motion trajectory of the user, with the teachings of Zhou of wherein having the performing the object detection further comprising: performing sampling on the gesture recognition result image data through frame extraction, to obtain sampled gesture recognition result image data; inputting the sampled gesture recognition result image data into an object detection model.
Wherein, Tong’s method of having wherein the performing the object detection further comprising performing sampling on the gesture recognition result image data through frame extraction, to obtain sampled gesture recognition result image data; inputting the sampled gesture recognition result image data into an object detection model.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform gesture detection more efficiently and be reduced on computation load. Since both Tong and Zhou’s systems perform hand gesture detection and tracking. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Zhou’s system improves performing gesture detection more efficiently and be reduced on computation load (see Zhou’s Par. [0098]).
Regarding claim 18, Tong in view of Ette discloses the interaction processing apparatus according to claim 12.
However, Tong does not explicitly teach wherein the processor is further configured to: input the sampled gesture recognition result image data, to determine the hand shape change and the gesture motion trajectory of the user.
Ette teaches wherein the processor is further configured to: (Pars. [0125-0126] disclose “a frame object has been detected within a frame of the image data…the outline of the frame object is parameterized to produce an interval profile”; furthermore, Pars. [0129-0131] disclose “the interval profile can be used to identify a maximum corresponding to an extended finger…a gesture can already be detected on the basis of the interval profile…the interval profile can be compared with one or more reference profiles…one or more profiles associated with the gesture detected in the present case, to be updated and stored”) input the gesture recognition result image data (Par. [0070] discloses “the profile comparison is performed on the basis of a machine learning method…the classification and detection of a gesture on the basis of the trajectory, the interval profile and the profile comparison” indicating a neural network for gesture detection [object detection]), to determine the hand shape change and the gesture motion trajectory of the user (Par. [0035] discloses “changes on the basis of the shape of the frame objects of successive frames can be checked for plausibility”; Par. [0120] discloses “detect a gesture on the basis of an interval profile and/or a trajectory”; furthermore, Par. [0133] discloses “the curve shape can be analyzed and, for example, the interval between the maxima associated with the extended fingers can be evaluated”, wherein the detection of the hand and its gesture can be used to further identify fingers and their maxima as part of an object detection to determine interval profile of a complete hand gesture as taught by Ette).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong of having an interaction processing apparatus to perform instructions comprises: receive a dynamic image of a gesture move of a user; perform gesture recognition on the dynamic image to obtain gesture recognition result image data of the dynamic image; to determine a hand shape change and a gesture motion trajectory of the user; determine, based on the hand shape change and the gesture motion trajectory, a gesture corresponding to the hand shape change and the gesture motion trajectory and an instruction mapped to the gesture; and execute the instruction, with the teachings of Ette of having wherein the processor is further configured to: input the sampled gesture recognition result image data, to determine the hand shape change and the gesture motion trajectory of the user.
Wherein Tong’s apparatus having wherein the processor is further configured to: input the sampled gesture recognition result image data, to determine the hand shape change and the gesture motion trajectory of the user.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform reliable detection of a capture gesture and reduce computation power. Since both Tong and Ette’s systems perform hand gesture detection. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Ette’s system improves on performing on a reliable detection of a capture gesture and reduce computation power (see Ette’ Par. [0014]).
However, Tong in view of Ette does not explicitly teach to perform sampling on the gesture recognition result image data through frame extraction, to obtain sampled gesture recognition result image data; input the sampled gesture recognition result image data into an object detection model.
Zhou teaches to perform sampling on the gesture recognition result image data through frame extraction (Par. [0063] discloses “determining the target bounding box based on the initial bounding boxes, the hand gesture detection apparatus…perform down-sampling processing on the initial bounding boxes to obtain spare bounding boxes”, wherein, the bounding boxes are gesture recognition result image data, down-sampling processing is analogous to sampling as recited, wherein the down-sampling is performed on the bounding boxes being feature extraction from the frame hence, is analogous to “through frame extraction”), to obtain sampled gesture recognition result image data (Par. [0064] discloses “the hand gesture detection apparatus selects the target bounding box from the spare bounding boxes obtained by the down-sampling processing”); input the sampled gesture recognition result image data into an object detection model (Par. [0064] discloses “the hand gesture detection apparatus selects the target bounding box from the spare bounding boxes obtained by the down-sampling processing”; furthermore, Par. [0075] discloses “the hand to be detected by using the gesture estimation model to obtain the gesture detection result of the hand to be detected”, wherein the gesture estimation model for detection of the hand is analogous to the object detection model as claimed).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Tong in view of Ette of having an interaction processing apparatus to execute instructions comprises: input the sampled gesture recognition result image data, to determine the hand shape change and the gesture motion trajectory of the user, with the teachings of Zhou of wherein having performing sampling on the gesture recognition result image data through frame extraction, to obtain sampled gesture recognition result image data; inputting the sampled gesture recognition result image data into an object detection model.
Wherein, Tong’s apparatus of having wherein performing sampling on the gesture recognition result image data through frame extraction, to obtain sampled gesture recognition result image data; inputting the sampled gesture recognition result image data into an object detection model.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform gesture detection more efficiently and be reduced on computation load. Since both Tong and Zhou’s systems perform hand gesture detection and tracking. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Zhou’s system improves performing gesture detection more efficiently and be reduced on computation load (see Zhou’s Par. [0098]).
Claims 8-9 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Xiaofeng Tong et. al. (“US 2013/0278504 A1” hereinafter as “Tong”) in view of Bernd Ette et. al. (“US 2020/0143154 A1” hereinafter as “Ette”) and Raffi Bedikian et. al. (“US 2019/0155394 A1” hereinafter as “Bedikian”).
Regarding claim 8, Tong in view of Ette teaches the interaction processing method according to claim 1, Tong teaches wherein the determining, based on the hand shape change and the gesture motion trajectory, the gesture (Par. [0030] discloses “use those hand features and the motion trajectory information provided by hand tracking module to identify dynamic hand gestures”) corresponding to the hand shape change and the gesture motion trajectory (Par. [0030] discloses “use those hand features and the motion trajectory information provided by hand tracking module to identify dynamic hand gestures”; moreover, Par. [0039] discloses “…a dynamic hand gesture based on the identified hand posture in combination with hand motion trajectory information”) and the instruction mapped to the gesture (Par. [0042] discloses “…that gesture identification provide gesture commands in response to the recognition of dynamic hand gestures” indicating an instruction [gesture commands]; furthermore, Par. [0041] discloses “…compare the dynamic hang gesture’s hand posture and motion feature to a database of hand postures and motion features of predefined dynamic hand gestures. For example, FIG. 5 depicts various predefined dynamic hand gestures and corresponding UI commands” a database of hand gesture and corresponding commands indicating mapping of corresponding information).
However, Tong in view of Ette does not explicitly teach searching a pre-established gesture library to determine the gesture corresponding to the hand shape change and the gesture motion trajectory and the instruction mapped to the gesture, wherein the gesture library records an association relationship between a gesture identifier, a hand shape change and a gesture motion trajectory corresponding to a gesture, and an instruction mapped to a gesture.
Bedikian teaches searching a pre-established gesture library to determine the gesture (Par. [0016] discloses “comparing the primitives to one or more templates in a library of gesture templates may include disassembling at least a portion of a trajectory into a set of frequency components…and searching for the set of frequency components among the template(s) stored in the library”) corresponding to the hand shape change and the gesture motion trajectory (Par. [0016] discloses “comparing the primitives to one or more templates in a library of gesture templates may include disassembling at least a portion of a trajectory”; Par. [0023] discloses “identify shapes and positions of the at least one control object in the images…to estimate a trajectory of the least one control object” indicating the trajectory is also based on shape detected for the object overtime; Par. [0079] discloses “second indication (e.g., changing the shape…)…the indication of the degree of gesture completion”) and the instruction mapped to the gesture (Par. [0080] discloses “the gesture-recognition system detects and identifies the user’s gestures based on the shales and positions of the gesturing part of the user’s body…once the gesture is recognized and the instruction associated therewith is identified” indicating an instruction associated with the gesture to be identified), wherein the gesture library records an association relationship between a gesture identifier (Par. [0091] discloses “an instruction associated with the gesture is identified by comparing the detected gesture with gestures stored in a database. Then, the ratio of the user’s actual movement to a resulting virtual action…is determined based on the instruction” the ratio here is analogous to the identifier which is used to identify the instruction associated with the gesture), a hand shape change and a gesture motion trajectory corresponding to a gesture (Par. [0080] discloses “the gesture-recognition system detects and identifies the user’s gestures based on the shales and positions of the gesturing part of the user’s body…once the gesture is recognized and the instruction associated therewith is identified”; Par. [0016] discloses “comparing the primitives to one or more templates in a library of gesture templates may include disassembling at least a portion of a trajectory into a set of frequency components…and searching for the set of frequency components among the template(s) stored in the library”), and an instruction mapped to a gesture (Par. [0091] discloses “an instruction associated with the gesture is identified by comparing the detected gesture with gestures stored in a database. Then, the ratio of the user’s actual movement to a resulting virtual action…is determined based on the instruction”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Tong in view of Ette of having an interaction processing method, wherein the determining, based on the hand shape change and the gesture motion trajectory, the gesture corresponding to the hand shape change and the gesture motion trajectory and the instruction mapped to the gesture, with the teachings of Bedikian of wherein having searching a pre-established gesture library to determine the gesture corresponding to the hand shape change and the gesture motion trajectory and the instruction mapped to the gesture, wherein the gesture library records an association relationship between a gesture identifier, a hand shape change and a gesture motion trajectory corresponding to a gesture, and an instruction mapped to a gesture.
Wherein, Tong’s method of having wherein searching a pre-established gesture library to determine the gesture corresponding to the hand shape change and the gesture motion trajectory and the instruction mapped to the gesture, wherein the gesture library records an association relationship between a gesture identifier, a hand shape change and a gesture motion trajectory corresponding to a gesture, and an instruction mapped to a gesture.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform improvement on machine interface using gesture detection for instructions of using such interface efficiently. Since both Tong and Bedikian’s systems perform hand gesture detection and tracking. Wherein Tong’s system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Bedikian’s system on machine interface using gesture detection for instructions of using such interface efficiently (see Bedikian’s Par. [0006]).
Regarding claim 9, Tong in view of Ette and Bedikian teaches the interaction processing method according to claim 8, Tong teaches further comprising: receiving a requirement for a gesture Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture in response to depth images of gesture” indicating a dynamic image of gesture [depth images of gesture]); acquiring a dynamic image of a gesture to form a basic data set (Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture in response to depth images of gesture” indicating a dynamic image of gesture [depth images of gesture being analogous to the recited basic data set]); performing gesture recognition on the basic data set (Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture” which is analogous to the recited gesture recognition [gesture recognition engine]) to obtain gesture recognition result image data of the gesture (Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture in response to depth images of gesture” wherein the dynamic hand gesture is analogous to the recited gesture recognition result image data; moreover, Par. [0020] discloses “…dynamic hand gesture including one or more hand postures combined with various hand motion trajectories”).
However, Tong in view of Bedikian does not explicitly teach performing object detection based on the gesture recognition result image data of the gesture, to obtain a hand shape change and a gesture motion trajectory of the gesture
Ette teaches performing object detection based on the gesture recognition result image data of the gesture (Pars. [0125-0126] disclose “a frame object has been detected within a frame of the image data…the outline of the frame object is parameterized to produce an interval profile”; furthermore, Pars. [0129-0131] disclose “the interval profile can be used to identify a maximum corresponding to an extended finger…a gesture can already be detected on the basis of the interval profile…the interval profile can be compared with one or more reference profiles…one or more profiles associated with the gesture detected in the present case, to be updated and stored”), to obtain a hand shape change and a gesture motion trajectory of the gesture (Par. [0035] discloses “changes on the basis of the shape of the frame objects of successive frames can be checked for plausibility”; Par. [0120] discloses “detect a gesture on the basis of an interval profile and/or a trajectory”; furthermore, Par. [0133] discloses “the curve shape can be analyzed and, for example, the interval between the maxima associated with the extended fingers can be evaluated” wherein, the detection of the hand and its gesture can be used to further identify fingers and their maxima as part of an object detection to determine interval profile of a complete hand gesture as taught by Ette).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong of having an interaction processing method, comprising: receiving a dynamic image of a gesture move of a user; performing gesture recognition on the dynamic image to obtain gesture recognition result image data of the dynamic image; to determine a hand shape change and a gesture motion trajectory of the user; determining, based on the hand shape change and the gesture motion trajectory, a gesture corresponding to the hand shape change and the gesture motion trajectory and an instruction mapped to the gesture; and executing the instruction, with the teachings of Ette of having wherein performing object detection based on the gesture recognition result image data of the gesture, to obtain a hand shape change and a gesture motion trajectory of the gesture.
Wherein Tong’s method having performing object detection based on the gesture recognition result image data of the gesture, to obtain a hand shape change and a gesture motion trajectory of the gesture.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform reliable detection of a capture gesture and be reduced on computation power. Since both Tong and Ette’s systems perform hand gesture detection. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Ette’s system improves performing on a reliable detection of a capture gesture and be reduced on computation power (see Ette’ Par. [0014]).
However, Tong in view of Ette does not explicit teach the gesture being a customized gesture from the user, to determine an identifier of the gesture and an instruction mapped to the gesture; and storing the identifier of the gesture, the hand shape change and the gesture motion trajectory of the gesture, and the instruction mapped to the gesture into the gesture library.
Bedikian teaches the gesture being a customized gesture from the user (Par. [0070] discloses “a new filter is generated or initiated every time new gestures are detected”; Par. [0071] discloses “map gestures onto the control inputs available with a computer mouse…free-space recognition in accordance herewith is not limited to traditional user-input actions, but facilities defining entirely new and distinct action” indicating facilities [users] can define new actions to customize the gesture for the processing hereby), to determine an identifier of the gesture and an instruction mapped to the gesture (Par. [0091] discloses “an instruction associated with the gesture is identified by comparing the detected gesture with gestures stored in a database. Then, the ratio of the user’s actual movement to a resulting virtual action…is determined based on the instruction” the ratio here is analogous to the identifier which is used to identify the instruction associated with the gesture); and storing the identifier of the gesture, the hand shape change and the gesture motion trajectory of the gesture (Par. [0080] discloses “the gesture-recognition system detects and identifies the user’s gestures based on the shales and positions of the gesturing part of the user’s body…once the gesture is recognized and the instruction associated therewith is identified”; Par. [0016] discloses “comparing the primitives to one or more templates in a library of gesture templates may include disassembling at least a portion of a trajectory into a set of frequency components…and searching for the set of frequency components among the template(s) stored in the library”), and the instruction mapped to the gesture into the gesture library (Par. [0091] discloses “an instruction associated with the gesture is identified by comparing the detected gesture with gestures stored in a database. Then, the ratio of the user’s actual movement to a resulting virtual action…is determined based on the instruction”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Tong in view of Ette of having an interaction processing method, wherein the determining, based on the hand shape change and the gesture motion trajectory, the gesture corresponding to the hand shape change and the gesture motion trajectory and the instruction mapped to the gesture, with the teachings of Bedikian of wherein having the gesture being a customized gesture from the user, to determine an identifier of the gesture and an instruction mapped to the gesture; and storing the identifier of the gesture, the hand shape change and the gesture motion trajectory of the gesture, and the instruction mapped to the gesture into the gesture library.
Wherein, Tong’s method of having the gesture being a customized gesture from the user, to determine an identifier of the gesture and an instruction mapped to the gesture; and storing the identifier of the gesture, the hand shape change and the gesture motion trajectory of the gesture, and the instruction mapped to the gesture into the gesture library.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform improvement on machine interface using gesture detection for instructions of using such interface efficiently. Since both Tong and Bedikian’s systems perform hand gesture detection. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Bedikian’s system improves on machine interface using gesture detection for instructions of using such interface efficiently (see Bedikian’s Par. [0006]).
Regarding claim 19, Tong in view of Ette teaches the interaction processing apparatus according to claim 12.
However, Tong in view of Ette does not explicitly teach wherein the processor is configured to: search a pre-established gesture library to determine the gesture corresponding to the hand shape change and the gesture motion trajectory and the instruction mapped to the gesture, wherein the gesture library records an association relationship between a gesture identifier, a hand shape change and a gesture motion trajectory corresponding to a gesture, and an instruction mapped to a gesture.
Bedikian teaches wherein the processor is configured to: search a pre-established gesture library to determine the gesture (Par. [0016] discloses “comparing the primitives to one or more templates in a library of gesture templates may include disassembling at least a portion of a trajectory into a set of frequency components…and searching for the set of frequency components among the template(s) stored in the library”) corresponding to the hand shape change and the gesture motion trajectory (Par. [0016] discloses “comparing the primitives to one or more templates in a library of gesture templates may include disassembling at least a portion of a trajectory”; Par. [0023] discloses “identify shapes and positions of the at least one control object in the images…to estimate a trajectory of the least one control object” indicating the trajectory is also based on shape detected for the object overtime; Par. [0079] discloses “second indication (e.g., changing the shape…)…the indication of the degree of gesture completion”) and the instruction mapped to the gesture (Par. [0080] discloses “the gesture-recognition system detects and identifies the user’s gestures based on the shales and positions of the gesturing part of the user’s body…once the gesture is recognized and the instruction associated therewith is identified” indicating an instruction associated with the gesture to be identified), wherein the gesture library records an association relationship between a gesture identifier (Par. [0091] discloses “an instruction associated with the gesture is identified by comparing the detected gesture with gestures stored in a database. Then, the ratio of the user’s actual movement to a resulting virtual action…is determined based on the instruction” the ratio here is analogous to the identifier which is used to identify the instruction associated with the gesture), a hand shape change and a gesture motion trajectory corresponding to a gesture (Par. [0080] discloses “the gesture-recognition system detects and identifies the user’s gestures based on the shales and positions of the gesturing part of the user’s body…once the gesture is recognized and the instruction associated therewith is identified”; Par. [0016] discloses “comparing the primitives to one or more templates in a library of gesture templates may include disassembling at least a portion of a trajectory into a set of frequency components…and searching for the set of frequency components among the template(s) stored in the library”), and an instruction mapped to a gesture (Par. [0091] discloses “an instruction associated with the gesture is identified by comparing the detected gesture with gestures stored in a database. Then, the ratio of the user’s actual movement to a resulting virtual action…is determined based on the instruction”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Tong in view of Ette of having an interaction processing apparatus, wherein the determining, based on the hand shape change and the gesture motion trajectory, the gesture corresponding to the hand shape change and the gesture motion trajectory and the instruction mapped to the gesture, with the teachings of Bedikian of wherein having the processor is further configured to perform: search a pre-established gesture library to determine the gesture corresponding to the hand shape change and the gesture motion trajectory and the instruction mapped to the gesture, wherein the gesture library records an association relationship between a gesture identifier, a hand shape change and a gesture motion trajectory corresponding to a gesture, and an instruction mapped to a gesture.
Wherein, Tong’s apparatus of having the processor is further configured to search a pre-established gesture library to determine the gesture corresponding to the hand shape change and the gesture motion trajectory and the instruction mapped to the gesture, wherein the gesture library records an association relationship between a gesture identifier, a hand shape change and a gesture motion trajectory corresponding to a gesture, and an instruction mapped to a gesture.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform improvement on machine interface using gesture detection for instructions of using such interface efficiently. Since both Tong and Bedikian’s systems perform hand gesture detection. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Bedikian’s system improves on machine interface using gesture detection for instructions of using such interface efficiently (see Bedikian’s Par. [0006]).
Regarding claim 20, Tong in view of Ette and Bedikian teaches the interaction processing apparatus according to claim 19, Tong teaches further wherein the processor is further configured to: receive a requirement for a gesture Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture in response to depth images of gesture” indicating a dynamic image of gesture [depth images of gesture]); acquire a dynamic image of a gesture to form a basic data set (Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture in response to depth images of gesture” indicating a dynamic image of gesture [depth images of gesture being analogous to the recited basic data set]); perform gesture recognition on the basic data set (Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture” which is analogous to the recited gesture recognition [gesture recognition engine]) to obtain gesture recognition result image data of the gesture (Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture in response to depth images of gesture” wherein the dynamic hand gesture is analogous to the recited gesture recognition result image data; moreover, Par. [0020] discloses “…dynamic hand gesture including one or more hand postures combined with various hand motion trajectories”).
However, Tong in view of Bedikian does not explicitly teach perform object detection based on the gesture recognition result image data of the gesture, to obtain a hand shape change and a gesture motion trajectory of the gesture
Ette teaches perform object detection based on the gesture recognition result image data of the gesture (Pars. [0125-0126] disclose “a frame object has been detected within a frame of the image data…the outline of the frame object is parameterized to produce an interval profile”; furthermore, Pars. [0129-0131] disclose “the interval profile can be used to identify a maximum corresponding to an extended finger…a gesture can already be detected on the basis of the interval profile…the interval profile can be compared with one or more reference profiles…one or more profiles associated with the gesture detected in the present case, to be updated and stored”), to obtain a hand shape change and a gesture motion trajectory of the gesture (Par. [0035] discloses “changes on the basis of the shape of the frame objects of successive frames can be checked for plausibility”; Par. [0120] discloses “detect a gesture on the basis of an interval profile and/or a trajectory”; furthermore, Par. [0133] discloses “the curve shape can be analyzed and, for example, the interval between the maxima associated with the extended fingers can be evaluated”, wherein the detection of the hand and its gesture can be used to further identify fingers and their maxima as part of an object detection to determine interval profile of a complete hand gesture as taught by Ette).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong of having an interaction processing apparatus to perform receiving a dynamic image of a gesture move of a user; performing gesture recognition on the dynamic image to obtain gesture recognition result image data of the dynamic image; to determine a hand shape change and a gesture motion trajectory of the user; determining, based on the hand shape change and the gesture motion trajectory, a gesture corresponding to the hand shape change and the gesture motion trajectory and an instruction mapped to the gesture; and executing the instruction, with the teachings of Ette of having wherein performing object detection based on the gesture recognition result image data of the gesture, to obtain a hand shape change and a gesture motion trajectory of the gesture.
Wherein having Tong’s apparatus wherein having performing object detection based on the gesture recognition result image data of the gesture, to obtain a hand shape change and a gesture motion trajectory of the gesture.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform reliable detection of a capture gesture and reduce computation power. Since both Tong and Ette’s systems perform hand gesture detection. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Ette’s system improves performing on a reliable detection of a capture gesture and reduce computation power (see Ette’ Par. [0014]).
However, Tong in view of Ette does not explicitly teach the gesture being a customized gesture from the user, to determine an identifier of the gesture and an instruction mapped to the gesture; and store the identifier of the gesture, the hand shape change and the gesture motion trajectory of the gesture, and the instruction mapped to the gesture into the gesture library.
Bedikian teaches the gesture being a customized gesture from the user (Par. [0070] discloses “a new filter is generated or initiated every time new gestures are detected”; Par. [0071] discloses “map gestures onto the control inputs available with a computer mouse…free-space recognition in accordance herewith is not limited to traditional user-input actions, but facilities defining entirely new and distinct action” indicating facilities [users] can define new actions to customize the gesture for the processing hereby), to determine an identifier of the gesture and an instruction mapped to the gesture (Par. [0091] discloses “an instruction associated with the gesture is identified by comparing the detected gesture with gestures stored in a database. Then, the ratio of the user’s actual movement to a resulting virtual action…is determined based on the instruction” the ratio here is analogous to the identifier which is used to identify the instruction associated with the gesture); and store the identifier of the gesture, the hand shape change and the gesture motion trajectory of the gesture (Par. [0080] discloses “the gesture-recognition system detects and identifies the user’s gestures based on the shales and positions of the gesturing part of the user’s body…once the gesture is recognized and the instruction associated therewith is identified”; Par. [0016] discloses “comparing the primitives to one or more templates in a library of gesture templates may include disassembling at least a portion of a trajectory into a set of frequency components…and searching for the set of frequency components among the template(s) stored in the library”), and the instruction mapped to the gesture into the gesture library (Par. [0091] discloses “an instruction associated with the gesture is identified by comparing the detected gesture with gestures stored in a database. Then, the ratio of the user’s actual movement to a resulting virtual action…is determined based on the instruction”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Tong in view of Ette of having an interaction processing apparatus wherein the determining, based on the hand shape change and the gesture motion trajectory, the gesture corresponding to the hand shape change and the gesture motion trajectory and the instruction mapped to the gesture, with the teachings of Bedikian of wherein having the gesture being a customized gesture from the user, to determine an identifier of the gesture and an instruction mapped to the gesture; and storing the identifier of the gesture, the hand shape change and the gesture motion trajectory of the gesture, and the instruction mapped to the gesture into the gesture library.
Wherein, Tong’s apparatus of having the gesture being a customized gesture from the user, to determine an identifier of the gesture and an instruction mapped to the gesture; and storing the identifier of the gesture, the hand shape change and the gesture motion trajectory of the gesture, and the instruction mapped to the gesture into the gesture library.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform improvement on machine interface using gesture detection for instructions of using such interface efficiently. Since both Tong and Bedikian’s systems perform hand gesture detection and tracking. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Bedikian’s system on machine interface using gesture detection for instructions of using such interface efficiently (see Bedikian’s Par. [0006]).
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Xiaofeng Tong et. al. (“US 2013/0278504 A1” hereinafter as “Tong”) in view of Bernd Ette et. al. (“US 2020/0143154 A1” hereinafter as “Ette”) further in view of Raffi Bedikian et. al. (“US 2019/0155394 A1” hereinafter as “Bedikian”) and Jonathan Marsden et. al. (“US 2023/0214458 A1” hereinafter as “Marsden”). and Christopher Allen Ingrassia et. al. (“US 2018/0053326 A1” hereinafter as “Ingrassia”).
Regarding claim 10, Tong in view of Ette and Bedikian teaches the interaction processing method according to claim 9, Tong teaches wherein the acquiring the dynamic image of the gesture to form the basic data set comprises (Par. [0016] discloses “…gesture recognition engine may be used to recognize a dynamic hand gesture in response to depth images of gesture” indicating a dynamic image of gesture [depth images of gesture being analogous to the recited basic data set]).
However, Tong in view of Ette does not explicitly teach the gesture being a customized gesture.
Bedikian teaches the gesture being a customized gesture (Par. [0070] discloses “a new filter is generated or initiated every time new gestures are detected”; Par. [0071] discloses “map gestures onto the control inputs available with a computer mouse…free-space recognition in accordance herewith is not limited to traditional user-input actions, but facilities defining entirely new and distinct action” indicating facilities [users] can define new actions to customize the gesture for the processing hereby).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Tong in view of Ette of having an interaction processing method, wherein the determining, based on the hand shape change and the gesture motion trajectory, the gesture corresponding to the hand shape change and the gesture motion trajectory and the instruction mapped to the gesture, with the teachings of Bedikian of wherein having the gesture being a customized gesture.
Wherein, Tong’s method of having wherein the gesture being a customized gesture.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform improvement on machine interface using gesture detection for instructions of using such interface efficiently. Since both Tong and Bedikian’s systems perform hand gesture detection and tracking. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Bedikian’s system improves on machine interface using gesture detection for instructions of using such interface efficiently (see Bedikian’s Par. [0006]).
However, Tong in view of Ette and Bedikian does not explicitly teach acquiring dynamic images of the gesture multiple times, with the dynamic image acquired each time forming a temporal image set.
Marsden teaches acquiring dynamic images (Par. [0333] discloses “the method further includes generating between 100,000 and 1 billion hand position-hand anatomy background simulations” wherein, the simulation is analogous to a dynamic image, a plurality of simulations includes a plurality of dynamic images), with the dynamic image acquired each time forming a temporal image set (Par. [0329] discloses “utilize a combination of a current simulated hand position image and a series of prior estimated hand position parameters temporally linked in previous frames to generate a current set of hand position parameters” indicating the simulation is used including temporally linked set).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette and Bedikian of having an interaction processing method, wherein the inputting the processed dynamic image into the gesture recognition model to obtain the gesture recognition result image data, with the teachings of Marsden of having wherein acquiring dynamic images of the gesture multiple times, with the dynamic image acquired each time forming a temporal image set.
Wherein, Tong’s method of having wherein acquiring dynamic images of the gesture multiple times, with the dynamic image acquired each time forming a temporal image set.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to improve gesture recognition using hand detection and hand pose estimation simulation for training of neural networks. Since both Tong and Marsden’s systems perform hand gesture detection. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with devices (see Tong’s Par. [0001]), and Marsden’s system improves gesture recognition using hand detection and hand pose estimation simulation for training of neural networks (see Marsden’s Par. [0078]).
However, Tong in view of Ette further in view of Bedikian and Marsden does not explicitly teach calculating an intersection of a plurality of temporal image sets to obtain the basic data set.
Ingrassia teaches calculating an intersection of a plurality of temporal image sets to obtain the basic data set (Par. [0027] discloses “if two or more images for two or more different data sets…are added, the results could be an image illustrating where those two data sets intersected each other and the color would indicate the proximity and magnitude…the map for intersection analysis can then be created”; furthermore, Par. [0024] discloses “data that changes in a temporal fashion. The attributed features…can have a temporal data aspect to it”; the map resulted from the intersection analysis is analogous to the recited basic data set).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teaches of Tong in view of Ette further in view of Bedikian and Marsden of having an interaction processing method, wherein the inputting the processed dynamic image into the gesture recognition model to obtain the gesture recognition result image data, wherein acquiring dynamic images of the gesture multiple times, with the dynamic image acquired each time forming a temporal image set, with the teachings of Ingrassia wherein having calculating an intersection of a plurality of temporal image sets to obtain the basic data set.
Wherein, Tong’s method of having wherein calculating an intersection of a plurality of temporal image sets to obtain the basic data set.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to allow data to be stored and queried and easy to be managed and pushing the most relevant and accurate data for processing. Since both Tong and Ingrassia’s systems perform data mapping and analysis. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Ingrassia’s system improves data management by allowing data to be stored and queried and easy to be managed and pushing the most relevant and accurate data for processing (see Ingrassia’s Par. [0012]).
Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Xiaofeng Tong et. al. (“US 2013/0278504 A1” hereinafter as “Tong”) in view of Bernd Ette et. al. (“US 2020/0143154 A1” hereinafter as “Ette”) and Raffi Bedikian et. al. (“US 2019/0155394 A1” hereinafter as “Bedikian”) and Jonathon Marsden et. al. (“US 2023/0214458 A1” hereinafter as “Marsden”).
Regarding claim 11, Tong in view of Ette and Bedikian teaches the interaction processing method according to claim 8, Tong teaches further comprising: receiving a rule for a gesture from the user (Par. [0025] discloses “a user may push his/her hand forward with closed posture to re-initiate hand tracking using hand detection module to detect the user’s hand” indicating a rule [initiating hand detection]); determining, according to the rule (Par. [0025] discloses “a user may push his/her hand forward with closed posture to re-initiate hand tracking using hand detection module to detect the user’s hand” indicating a rule [initiating hand detection], initiating the hand detection), an identifier of the gesture (Par. [0028] discloses “hand tracking module may include skin color identification ion codes (or instruction sets) that are generally operable to distinguish…”, wherein the identification code is analogous to an identifier of the gesture as claimed), a definition of the gesture (Par. [0039] discloses “gesture identification module may then determine, based on the hand posture and trajectory, which type of pre-defined dynamic hand gesture the identified dynamic hand gesture corresponds to” indicating a definition of the gesture [pre-defined type]), and an instruction mapped to the gesture (Par. [0044] discloses “the type of dynamic hand gesture identified. This, in turn may be used by gesture command module to generate UI-command”, wherein the command is analogous to the instruction being associated with the gesture).
However, Tong in view of Ette does not explicitly teach and storing the identifier of the gesture, the hand shape change and the gesture motion trajectory of the gesture, and the instruction mapped to the gesture into the gesture library.
Bedikian teaches and storing the identifier of the gesture, the hand shape change and the gesture motion trajectory of the gesture (Par. [0080] discloses “the gesture-recognition system detects and identifies the user’s gestures based on the shales and positions of the gesturing part of the user’s body…once the gesture is recognized and the instruction associated therewith is identified”; Par. [0016] discloses “comparing the primitives to one or more templates in a library of gesture templates may include disassembling at least a portion of a trajectory into a set of frequency components…and searching for the set of frequency components among the template(s) stored in the library”), and the instruction mapped to the gesture into the gesture library (Par. [0091] discloses “an instruction associated with the gesture is identified by comparing the detected gesture with gestures stored in a database. Then, the ratio of the user’s actual movement to a resulting virtual action…is determined based on the instruction”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Tong in view of Ette of having an interaction processing method, wherein the determining, based on the hand shape change and the gesture motion trajectory, the gesture corresponding to the hand shape change and the gesture motion trajectory and the instruction mapped to the gesture, with the teachings of Bedikian of wherein having storing the identifier of the gesture, the hand shape change and the gesture motion trajectory of the gesture, and the instruction mapped to the gesture into the gesture library.
Wherein, Tong’s method of having wherein storing the identifier of the gesture, the hand shape change and the gesture motion trajectory of the gesture, and the instruction mapped to the gesture into the gesture library.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to perform improvement on machine interface using gesture detection for instructions of using such interface efficiently. Since both Tong and Bedikian’s systems perform hand gesture detection. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Bedikian’s system improves on machine interface using gesture detection for instructions of using such interface efficiently (see Bedikian’s Par. [0006]).
However, Tong in view of Ette and Bedikian does not explicitly teach simulating a hand shape change and a gesture motion trajectory of the gesture according to the definition of the gesture.
Marsden teaches simulating a hand shape change (Par. [0057] discloses “computer graphic simulator that prepares sample simulated hand positions of gesture sequences for training of neural networks”; moreover, Par. [0260] discloses “various simulation parameters of the selected hand such as position, shape, size, etc. are adjusted by moving or reshaping the hand”) and a gesture motion trajectory of the gesture (Par. [0348] discloses “hand-heuristic analysis determines whether the detected hand is a right hand or a left hand based on an estimated trajectory of the particular hand region”) according to the definition of the gesture (Par. [0257] discloses “rendering type of the simulated hand is defined by the rendering attributes”).
Therefore, it would have been obvious to one or ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Tong in view of Ette and Bedikian of having an interaction processing method, wherein the determining, based on the hand shape change and the gesture motion trajectory, the gesture corresponding to the hand shape change and the gesture motion trajectory and the instruction mapped to the gesture, wherein having storing the identifier of the gesture, the hand shape change and the gesture motion trajectory of the gesture, and the instruction mapped to the gesture into the gesture library, with the teachings of Marsden of having wherein simulating a hand shape change and a gesture motion trajectory of the gesture according to the definition of the gesture.
Wherein, Tong’s method of wherein simulating a hand shape change and a gesture motion trajectory of the gesture according to the definition of the gesture.
The motivation behind the modification would have been to perform gesture detection for improving user interfaces and user interaction with hand gesture interactive devices, and further to improve gesture recognition using hand detection and hand pose estimation simulation for training of neural networks. Since both Tong and Marsden’s systems perform hand gesture detection and action processing. Wherein Tong system improves gesture detection for improving user interfaces and user interaction with hand gesture interactive devices (see Tong’s Par. [0001]), and Marsden’s system improves gesture recognition using hand detection and hand pose estimation simulation for training of neural networks (see Marsden’s Par. [0078]).
Pertinent Prior Art(s)
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Brown, Colin J. et. al., “US 2025/0335766 A1”, discloses an activity classifier system and method that classifies human activities using 2D skeleton data. The system includes a skeleton preprocessor that transforms the 2D skeleton data into transformed skeleton data, the transformed skeleton data comprising scaled, relative joint positions and relative joint velocities. The system also includes a gesture classifier comprising a first recurrent neural network that receives the transformed skeleton data, and is trained to identify the most probable of a plurality of gestures. The system also has an action classifier comprising a second recurrent neural network that receives information from the first recurrent neural networks and is trained to identify the most probable of a plurality of actions.
HIROSE, Mitsunobu et. al., “US 2022/0044008 A1”, discloses a system, a method, and a program that easily verify that a certificate with a photograph belongs to a user. The system that verifies that a certificate with a photograph belongs to a user acquires a first image containing the certificate with the photograph of a user, judges the validity of the first image, acquires a second image containing the user and the certificate with a photograph that corresponds to the first image that has validity, judges if the user's face and the photograph of the certificate in the second image match, and certifies that the certificate with a photograph belongs to the user if the face and the photograph of the certificate match.
Zhang, Lidan et. al., “US 2023/0410487 A1”, discloses performing online learning for a model to detect unseen actions in an action recognition system is disclosed. The method includes extracting semantic features in a semantic domain from semantic action labels, transforming the semantic features from the semantic domain into mixed features in a mixed domain, and storing the mixed features in a feature database. The method further includes extracting visual features in a visual domain from a video stream and determining if the visual features indicate an unseen action in the video stream. If no unseen action is determined, applying an offline classification model to the visual features to identify seen actions, assigning identifiers to the identified seen actions, transforming the visual features from the visual domain into mixed features in the mixed domain, and storing the mixed features and seen action identifiers in the feature database. If an unseen action is determined, transforming the visual features from the visual domain into mixed features in the mixed domain, applying a continual learner model to mixed features from the feature database to identify unseen actions in the video stream, assigning identifiers to the identified unseen actions, and storing the unseen action identifiers in the feature database.
Liu, Yen Ting et. al., “US 2023/0145728 A1”, discloses a method and a system for detecting a hand gesture, and a computer readable storage medium. The method includes: determining whether information of a hand is enough for identifying a hand gesture of the hand; in response to determining that the information of the hand is enough for identifying the hand gesture, identifying the hand gesture, receiving first hand gesture information from at least one external gesture information provider, and correcting the hand gesture based on the first hand gesture information; in response to determining that the information of the hand is not enough for identifying the hand gesture, receiving second hand gesture information from the at least one external gesture information provider, obtaining a predicted hand gesture, and obtaining the hand gesture based on the predicted hand gesture and the second hand gesture information.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHUONG HAU CAI whose telephone number is (571)272-9424. The examiner can normally be reached M-F 8:30 am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PHUONG HAU CAI/Examiner, Art Unit 2673
/CHINEYERE WILLS-BURNS/Supervisory Patent Examiner, Art Unit 2673