DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of the Claims
Claims 1-20 are rejected on the grounds of non-statutory double patenting.
Claims 1-2, 4, 7-9, 11, 14-16, 18, and 20 are rejected under 35 U.S.C. 102(a)(1).
Claims 3, 5-6, 10, 12-13, 17, and 19 are rejected under 35 U.S.C. 103.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 1-20 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1 and 4-7 of U.S. Patent No. US 12,282,607 B2 (hereinfacter ‘607). Although the claims at issue are not identical, they are not patentably distinct from each other because the claims of the instant application are encompassed within the claims of ‘607, as can be seen from the comparisons provided in the below table.
Instant Application
Claims of ‘607
Comments
(Claim 1) A machine-implemented method comprising: determining, by an AR system, a language model based on registration data of an AR application, the registration data including an ID field identifying the AR application, a language field identifying the language model, and one or more symbol fields indicating respective one or more symbols to be routed to the AR application;
(claim 1) determining, by one or more processors, a language model based on component registration data of the AR application component, the component registration data including a component ID field identifying the AR application component, a language field identifying the language model, and one or more symbol fields indicating respective one or more symbols to be routed to the AR application component
Substantially similar.
The same comparison applies to claims 8 and 15.
detecting, by the AR system, using one or more cameras and the language model, one or more input symbols corresponding to one or more gestures being made by a user of the AR application;
detecting, by the one or more processors, using the one or more cameras and the language model, one or more input symbols corresponding to fingerspelling signs being made by the user
Substantially similar.
routing, by the AR system, the one or more input symbols to the AR application using the one or more input symbols and the one or more symbol fields of the registration data;
routing, by the one or more processors, the one or more input symbols to the AR application component using the one or more input symbols and the one or more symbol fields of the component registration data;
Substantially similar.
and providing, by the AR application, a user interface using the one or more input symbols.
generating, by the one or more processors, entered text data based on the one or more input symbols; and providing, by the one or more processors, text in the text scene component based on the entered text data.
Substantially similar.
(Claim 2) wherein routing the one or more input symbols to the AR application further comprises: classifying symbol input event data of the one or more input symbols as directed input event data based on the symbol input event data and the registration data.
(claim 5) wherein routing the entered text data to the AR application component further comprises: classifying symbol input event data of the one or more input symbols as directed input event data based on the symbol input event data and the component registration data.
Substantially similar.
The same comparison applies to claims 9 and 16.
(Claim 3) wherein the language model is trained on a corpus of text related to a context of the user interface.
(claim 4) wherein the language model comprises a hidden Markov model trained on a corpus of text related to a context of the text scene component.
Substantially similar.
The same comparison applies to claims 10 and 17.
(Claim 4) wherein the user interface includes a text component.
(claim 1) … providing, by an Augmented Reality (AR) system, to a user of the AR system, a text scene component in a user interface
Substantially similar.
The same comparison applies to claims 11 and 18.
(Claim 5) wherein the one or more input symbols include one or more editing commands, and wherein providing the user interface using the one or more input symbols further includes performing an editing function on text data of the text component.
(claim 6) wherein the one or more input symbols include one or more editing commands, and wherein generating the entered text data based on the one or more input symbols further includes performing an editing function on the entered text data
Substantially similar.
The same comparison applies to claims 12 and 19.
(Claim 6) further comprising: determining, by the AR system, that an apparent location of the text component is correlated to a location in a real-world scene environment of a hand of the user while the user is making a gesture.
(claim 1) …the text scene component having an apparent location in a real-world scene environment… determining, by the one or more processors, that the apparent location of the text scene component is correlated to a location in the real-world scene environment of a hand of the user while the user is making the start text entry gesture
Substantially similar.
The same comparison applies to claim 13.
(Claim 7) wherein the AR system comprises a head-worn device.
(claim 7) wherein the AR system comprises a head-worn device.
Identical.
The same comparison applies to claims 14 and 20.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-2, 4, 7-9, 11, 14-16, 18, and 20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by BROWY (US 2018/0075659 A1).
Regarding Claim 1, BROWY teaches a machine-implemented method comprising: determining, by an AR system (¶ 32-33, 36: A wearable AR or mixed reality system includes sign language recognition), a language model based on registration data of an AR application, (“For example, the wearable system 200 can include sign-language-to-text functionality implemented on the HMD. As one example, the wearable system can store a sign language dictionary in the local data module 260 or the remote data repository 280. The wearable system can accordingly access the sign language dictionary to translate a detected gesture into text.” Paragraph 153. A language model is determined for converting captured sign language into text. See Paragraph 152, which further discusses that this conversion is done using deep learning or hidden Markov models, and Paragraph 158, which discusses that multiple sign language models for different sign languages are stored by the system. Sign language recognition can be reconfigured, which reads on determining the language model based on “registration data”.)
the registration data including an ID field identifying the AR application, (“the wearable system may identify a particular UI. The type of UI may be predetermined by the user. The wearable system may identify that a particular UI needs to be populated based on a user input” Paragraph 128. The wearable system identifies a particular or type of UI, which is registering of an identification of the AR application.)
a language field identifying the language model, (“the wearable system 200 may also be configured to recognize a plurality of sign languages, e.g., ASL, British Sign Language, Chinese Sign Language, Dogon Sign Language, etc. In some implementations, the wearable system 200 supports reconfiguration of sign language recognition, e.g., based on location information of the sensory eyewear system… a user can select a foreign sign language and the wearable system can present the meaning of gestures in the foreign sign language as text to the user of the wearable system.” Paragraph 158: A language field is identified when configuring the wearable device, either automatically or by the user.)
and one or more symbol fields indicating respective one or more symbols to be routed to the AR application; (Paragraphs 176-178, Fig. 3A with context from Paragraphs 152-153: A dictionary is maintained comprising a plurality of sign language symbols and their corresponding text. These entries in the dictionary read on “symbol fields”. After recognition of a sign being gestured by a user, the symbol would be routed to the AR application and displayed in some way, such as by a text component as shown in Fig. 3D.)
detecting, by the AR system, using one or more cameras and the language model, one or more input symbols corresponding to one or more gestures being made by a user of the AR application… and providing, by the AR application, a user interface using the one or more input symbols. (“the wearable system 200 can interpret sign language by, for example, detecting gestures that may constitute sign language, translating the sign language to another language (e.g., another sign language or a spoken language), and presenting the translated information to a user of a wearable device” Paragraph 139. “the user of the wearable system 200 may also communicate with sign language. In this case, the wearable system can capture the user's own sign language gestures (from a first-person point of view) by the outward-facing imaging system 464. The wearable system can convert the sign language to a target speech which may be expressed in the format of text, audio, images, etc” Paragraph 146. “The wearable system 200 can convert the captured sign language to text which can be presented to a user or translated into another language. Conversion of sign language to text can be performed using algorithms such as deep learning (which may utilize a deep neural network), hidden Markov model, dynamic programming matching, etc.” Paragraph 152. “The converted text or auxiliary information may be presented in a variety of ways. In one example, the wearable system 200 can place converted text or auxiliary information in text bubbles” Paragraph 168. Also see Paragraphs 178, 180, 191, 194, 197 and Figure 14B, which further discuss detecting sign language of the user wearing the AR device and converting the user’s sign language to text to be displayed on a component of one or more AR devices.)
routing, by the AR system, the one or more input symbols to the AR application using the one or more input symbols and the one or more symbol fields of the registration data; (“The wearable device 1380 can present a virtual user interface 1382 to show the input 1384a captured by the wearable device, the translation 1384b corresponding to the input 1384a (e.g., “Is there a pharmacy nearby? I'm feeling unwell.”).” Paragraph 186. The one or more input symbols are routed to a component of the AR system as shown in Figure 13D. In context with Paragraphs 120 and 152-153, the component would be configured in a language that the user of the AR system understands using a dictionary and machine learning models (i.e. language model) corresponding to a language setting selected by a user. Therefore, the routing is done using the one or more input symbols and “component registration data”. It is further obvious this would occur whether the user is observing the signer with the AR device or if the user is the signer themself since both embodiments are taught.)
Claim 8 is directed to a machine and claim 15 is directed to a machine-readable medium but they otherwise recite the same limitations as claim 1. Claim 8 and Claim 15 are therefore rejected using the same reasoning discussed above.
Regarding Claim 2, BROWY further teaches wherein routing the one or more input symbols to the AR application further comprises: classifying symbol input event data of the one or more input symbols as directed input event data based on the symbol input event data (BROWY, “The wearable device 1380 can observe (e.g., via the outward-facing imaging system 464) the hand gestures by the user 1394 as shown in FIG. 13D. The wearable device 1380 can automatically (e.g., using object recognizers 708) detect that the hand gestures as shown are an expression in a sign language, recognize the meaning associated with the hand gestures, and provide the translation of the hand gestures in a target language (e.g., English) which the user 1392 understands… The wearable system can use various object recognizers to detect the presence of hand gestures. For example, the wearable system may find a sequence of hand gestures may constitute a phrase or a sentence in a sign language… At block 1452, the example system can determine whether the detected sign language is the user's own… The wearable device of the signer can convert the signer's own sign language to text. The wearable device can transmit the converted text to a remote system viewable by the second person.” Paragraph 186, 189, 197, 206. Object recognizers classify hand gestures made by a user as sign language. The sign language is to be converted into text to be displayed in an AR application component, such as an example shown in Figure 3D. The AR system also determines if it is the user that is signing for the purpose of inputting or communicating text. Since a determination is made of which user is signing and a determination is made that the hand gestures are to be interpreted as sign language gestures, then a determination is being made that the input symbols are “directed input event data” based on the recognized input data. The examiner is interpreting “directed input event data” as input event data corresponding to sign language symbols that are intended to be input by the user of the device.)
and the registration data. (BROWY, Paragraph 120, which teaches that the UI of the AR system is set by the user to detect and convert sign language according to a list designated by the user. This is equivalent to “component registration data”. If a component is registered to receive text, it would have been obvious to classify the sign language input symbols as “directed data” to be displayed by the component, such as the component shown in Figure 13D.)
Claim 9 and Claim 16 recite the same limitations as claim 2 and are rejected for the same reasoning discussed above.
Regarding Claim 4, BROWY further teaches wherein the user interface includes a text component. (BROWY, ¶ 120: the UI includes data entry fields, which are text components. Also see Fig. 13D and ¶ 186, which describe another embodiment of a text component that is displayed with the input from the sign language recognition.)
Claim 11 and Claim 18 recite the same limitations as claim 4 and are rejected for the same reasoning discussed above.
Regarding Claim 7, BROWY further teaches wherein the AR system comprises a head-worn device. (BROWY, Figure 2A, illustrates that the AR system comprises a head-worn device.)
Claim 14 and Claim 20 recite the same limitations as claim 7 and are rejected for the same reasoning discussed above.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 3, 10, and 17 are rejected under 35 U.S.C. 103 as being unpatentable over BROWY (US 2018/0075659 A1) in view of GALOR (US 20120078614 A1).
Regarding Claim 3, BROWY teaches all the limiations of claim 1, on which claim 3 depends.
BROWY does not explicitly teach wherein the language model is trained on a corpus of text related to a context of the user interface.
However, GALOR, which is directed to text input into a 3D user interface, teaches wherein the language model is trained on a corpus of text related to a context of the user interface. “the language model may utilize an expected semantic domain. For example, the language model may select a response using a dictionary custom tailored to a question or a field type that 3D user interface 20 presents on display 28. In other words, the language model may utilize a custom dictionary specific to an application executing on computer 26… Examples of language models that can be implemented by computer 26 include a dictionary and statistical models including but not limited to a statistical dictionary, an n-gram model, a Markov model, and a dynamic Bayesian network” Paragraphs 50-51. A language model, which comprises a Markov model or Bayesian network, is tailored (i.e. trained) to the field of the application based on a context of the field or application. See Paragraph 19, which discusses that the language model is used to predict the input of the user using the camera of the device.)
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to modify the entering of text into an AR application component using detected sign language input by a user of the AR system taught by BROWY by training the language model to be based on the context of the input field as taught by GALOR since such an implementation would have enabled faster input of text since the output options would be narrowed down (See Paragraph 50). BROWY (Paragraph 152) also teaches that hidden Markov models can be trained using data containing known signs (supervised learning) in order to assist in classification and detection of gestures made by subsequent users. As taught by GALOR (Paragraph 21), “Utilizing a language model can provide a best guess of the user's intended input that enables the user to enter characters (i.e., via the virtual keyboard) [or via sign language detection in view of BROWY] more rapidly.”
Claim 10 and Claim 17 recite the same limitations as claim 3 and are rejected for the same reasoning discussed above.
Claims 5, 12, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over BROWY (US 2018/0075659 A1) in view of IVERS (US 2018/0114366 A1).
Regarding Claim 5, BROWY teaches all the limitations of claim 4, on which claim 5 depends.
BROWY does not teach wherein the one or more input symbols include one or more editing commands, and wherein providing the user interface using the one or more input symbols further includes performing an editing function on text data of the text component.
However, IVERS, which is directed to text editing in an augmented reality interface, teaches wherein the one or more input symbols include one or more editing commands, and wherein providing the user interface using the one or more input symbols further includes performing an editing function on text data of the text component. (“Gestures may be user-defined and configured. For example, a “double-pointing” motion may indicate to begin translation services. As another example, a single hand motion with two fingers pointing at the object with text may indicate to begin a copy text function. After copying text, the user may access another application in AR, such as an Internet browser, and paste the text into a user interface text control.” Paragraph 0022. See Figure 5B, which shows a user gesture in an AR viewer that results in an editing function being performed on the text, including a copy and a translate operation.)
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to modify the AR system for text entry using fingerspelling taught by BROWY in view of GALOR by incorporating the user hand gestures that result in text editing functions taught by IVERS. Since the references are directed to text entry or manipulation using hand gestures in an AR environment, the combination would have yielded predictable results. As suggested by IVERS (Paragraph 0053), this would improve the user experience by allowing the user to more easily navigate between and copy and paste content in different views of the AR application.
Claim 12 and Claim 19 recite the same limitations as claim 5 and are rejected for the same reasoning discussed above.
Claims 6 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over BROWY (US 2018/0075659 A1) in view of FURTWANGLER (US 2019/0339837 A1) and further in view of ZHOU (US 9,207,852 B1).
Regarding Claim 6, BROWY teaches all the limitations of claim 4, on which claim 6 depends.
BROWY does not teach further comprising: determining, by the AR system, that an apparent location of the text component is correlated to a location in a real-world scene environment of a hand of the user while the user is making a gesture.
However, FURTWANGLER, which is directed to a text entry process in an AR system, teaches further comprising: determining, by the AR system, that an apparent location of the text component is correlated to a location in a real-world scene environment of a hand of the user… (“As illustrated in FIG. 2B, the user 101 may select the media 216c by pointing the pointer 212 at a desired location and inputting an input into the virtual reality input device(s) 134.” Paragraph 47. “FIG. 6A illustrates the user 101 directing the virtual reality input device 134 towards an interactive field 604 (e.g., search box) of the application (e.g., Facebook) with a pointer 606 and pointer path 608. In particular embodiments, panel 602 may display any kind of application as described above. As illustrated in FIG. 6A, the user 101 may be able to “click” on the interactive field 604 of the application by pointing the pointer 606 at a desired location and inputting an input into the virtual reality input device 134 (e.g., clicking a button).” Paragraph 55. “the client system 130 may monitor the position of the pointer 906 on the panel 902 and determine when a gesture is made that encloses content (e.g., a loop, a circle, etc.) within a short time frame (e.g., 3 seconds). In particular embodiments, the user 101 may need to input an input into the virtual reality input device(s) 134 (e.g., click a button) to initiate the gesture.” Paragraph 62. A determination is made that the location pointed to by the hand of the user corresponds to a location of a UI element in the AR environment, including a text box.)
Furthermore, ZHOU, which is directed to recognizing hand gestures to manipulate text elements on a display, teaches the determination that the location of the text component is correlated to the location of the hand of the user while the user is making a gesture. (“as illustrated in FIG. 5(a) the text box 508 could have a highlight, bolded border, flashing element, or other indicia indicating that the text box corresponds to a current location of the user's hand based on the direction information 516 obtained by the camera 510 of the device. Such indicia can enable the user to know that the performance of a selection action will select that particular element” Col. 8:35-40. “In this particular example, however, the user is able to take advantage of the fact that the user's hand position is already being captured and analyzed in image information to utilize a specific gesture or more to provide the indication.” Col. 9:5-8. While the camera is determining that a user’s hand is positioned with respect to a particular element, such as a text box, the camera also determines if the user is making the start text entry gesture.)
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to modify the AR system for inputting text into a component of an application using fingerspelling taught by BROWY by detecting the location of the component and the user’s hand are correlated while the hand gesture for inputting text into the component is being performed as taught by the combination of FURTWANGLER and ZHOU. Since both references are directed to inputting text into a component of a display using gestures, the combination would have yielded predictable results and would have amounted to applying the teachings of FURTWANGLER and ZHOU to the AR sign language recognition device taught by BROWY. Since FURTWANGLER (Paragraph 45) teaches that the headset device includes cameras for overlaying virtual objects on a real-world scene, selecting and initiating inputting text data based on a camera-detected hand gesture would have been an obvious alternative input method that is taught by ZHOU. Furthermore, ZHOU (Col. 8:52-55) teaches the importance determining correlation between the hand gesture and the input intention in that it allows the user to indicate that they actually want to provide input at a particular location.
Claim 13 recites the same limitations as claim 6 and are rejected for the same reasoning discussed above.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure.
Shaffer (US 2020/0110927 A1) teaches imaging of a user using sign language, including facial expressions, where the method includes determining an ID of the user. (¶ 3-4)
Arana (US 2022/0358854 A1) teaches AR glasses that render the performance of signal language, including dynamic translation. (¶ 43)
Adamo (US 2006/0087510 A1) teaches configuration of sign language symbols within a database. (Fig. 4, ¶ 149)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to RAMI RAFAT OKASHA whose telephone number is (571)272-0675. The examiner can normally be reached M-F 10-6 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, SCOTT BADERMAN can be reached at (571) 272-3644. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RAMI R OKASHA/Primary Examiner, Art Unit 2118