Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Priority
No foreign or domestic priority is claimed. The effective filing date of U.S. Application No. 19/050,581 is 02/11/2025.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title.
Claims 12-16 are rejected under 35 U.S.C. § 101 because the claimed “computer-readable storage medium” encompasses non-statutory transitory signal embodiments. The Specification states that data and message structures may be transmitted via “a signal on a communications link” and describes computer-readable media as comprising both computer-readable storage media, e.g., “non-transitory” media, and computer-readable transmission media. Spec., ¶ [0082]. Thus, claim 12 does not limit the claimed medium to only non-transitory embodiments, and its broadest reasonable interpretation encompasses a transitory propagating signal, which is not a process, machine, manufacture, or composition of matter under 35 U.S.C. § 101.
Status of Claims
Claims 1–20 are pending in the application. Claims 1-5, 7-9, 11-19 are rejected.
Claims 6, 10, 20 are objected to.
Allowable Subject Matter
Claims 6, 10, 20 are objected to as being dependent upon a rejected base claim(s), but would be allowable if rewritten in independent form including all of the limitations of the base claim(s) and any intervening claim(s).
Overview of Grounds of Rejection
Ground of Rejection
Claim(s)
Statute(s) (e.g., § 102, § 103)
Reference(s)
12-16
§ 101
Ground 1
1–5, 7–9, 12–19
§ 103
Paul et al. (US20220337741A1) in view of Swaminathan et al. (US20180253145A1)
Ground 2
11
§ 103
Paul et al. (US20220337741A1) in view of Swaminathan et al. (US20180253145A1), and further in view of Parkinson et al. (US20130289971A1)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
(Please see the cited paragraphs, sections, pages, or surrounding text in the references for the paraphrased content.)
Ground of Rejection 1
Claims 1, 2, 3, 4, 5, 7, 8, 9, 12, 13, 14, 15, 16, 17, 18, 19 are rejected under 35 U.S.C. § 103 as being unpatentable over Paul et al. (US20220337741A1) in view of Swaminathan et al. (US20180253145A1).
As per Claim 1, Paul teaches the following portion of Claim 1, which recites: “A method for automatically translating text from a first language to a second language, the method comprising:”
Paul et al. teaches a system configured to “translate German words ... into English words”, corresponding to translation from a first language to a second language. Paul et al., ¶ [0518].
Paul alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Swaminathan, they collectively teach all of the limitation(s).
Paul and Swaminathan teach the following portion of Claim 1, which recites: “detecting, by an artificial reality system, one or more text segments, in the first language, in a field-of-view of the artificial reality system in a real-world environment;”
Paul et al. teaches a “live preview 1430 [that] is a representation of the FOV” containing “German words 1444 c 1-1444 n,” thereby detecting multiple first-language text portions in a camera field-of-view. Paul et al., ¶¶ [0523]-[0525].
Paul does not clearly teach the claimed artificial-reality aspect. Swaminathan et al. teaches that “road signs ... documents, and other real-world 2D and 3D objects can be captured with a front-facing camera in an AR view mode.” Swaminathan et al., ¶ [0039].
Paul teaches the following portion of Claim 1, which recites: “automatically translating the one or more text segments from the first language to the second language;”
Paul et al. teaches automatically, “without intervening user input and/or gestures,” displaying a plurality of translated-text indications, including translations of first and second text portions, and states that this includes “automatically translating the text of the representation of the field-of-view.” Paul et al., ¶ [0547].
Swaminathan teaches the following portion of Claim 1, which recites: “obtaining, by the artificial reality system, user gaze input;”
Swaminathan et al. teaches determining an “area of interest based on the user's eye gaze” using an eye-gaze processing algorithm. Swaminathan et al., ¶ [0077].
Swaminathan teaches the following portion of Claim 1, which recites: “determining an intent to interact with at least a portion of the one or more text segments, in the first language, based on the user gaze input;”
Swaminathan et al. teaches that “Eye gaze tracking can be used to evaluate a user's interest” and determines whether the gaze-defined area of interest is “lingering on or about the feature or the displayed object tag,” with corresponding AR information presented when the linger threshold is satisfied. Swaminathan et al., ¶¶ [0038], [0079]-[0080].
Paul teaches the following portion of Claim 1, which recites: “based on the determining the intent to interact with at least the portion of the one or more text segments, selecting, from the automatically translated one or more text segments, one or more translated text segments corresponding to at least the portion of the one or more text segments;”
Paul et al. teaches that, after displaying multiple translated portions, the system receives a request “to select a respective indication ... of the plurality of translated portions.” Paul et al., ¶ [0548]. Paul further teaches that selection of the first indication causes display of the translation of the first portion without the translation of the second portion. Paul et al., ¶ [0549]. Thus, Paul teaches selecting a corresponding translated portion from previously generated translated portions, while Swaminathan supplies the gaze-based intent used for that selection.
Paul teaches the following portion of Claim 1, which recites: “and rendering the at least the selected one or more translated text segments.”
Paul et al. teaches that selection of translated object “EGGS” causes the system to “display[] translation card 1470,” including the translated word “EGGS” in English. Paul et al., ¶¶ [0526]-[0527].
Before the effective filing date of the claimed invention, a person of ordinary skill in the art would have been motivated to modify Paul's camera-based automatic translation and translated-text selection system with Swaminathan's gaze-based AR interaction technique because Swaminathan teaches using eye gaze and gaze linger to identify a user's area of interest and selectively present corresponding AR information. Applying that known gaze-selection technique to Paul's already-translated text portions would improve hands-free interaction and reduce unnecessary display or selection of unrelated translations, yielding the predictable result of selecting and rendering the translated text corresponding to the user's gaze-indicated intent.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
As per Claim 2, Paul alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Swaminathan, they collectively teach all of the limitation(s).
Paul and Swaminathan teach Claim 2, which recites: “The method of claim 1, wherein the rendering the selected one or more translated text segments includes: positioning the selected one or more translated text segments locked at a location, in the real-world environment, corresponding to a determined location of the at least the portion of the one or more text segments in the first language.”
Paul et al. teaches positioning translated text “at a location corresponding to the original text that has been translated” and displaying a translated indication “on top of ... the first portion of the text,” thereby positioning the translation at the corresponding source-text location. Paul et al., ¶¶ [0547], [0563]. For the claimed “locked at a location, in the real-world environment”, Swaminathan et al. teaches an AR tracking algorithm such that “the displayed augmentation information follows (e.g., tracks) the target object”, with pose estimation and tracking used to maintain the rendered AR information relative to the target as the camera orientation changes. Swaminathan et al., ¶ [0069]. Thus, the combination teaches keeping Paul's corresponding translated text spatially associated with the real-world source-text location. The rationale and motivation to combine the references as set forth for claim 1 are incorporated herein by reference for the present claim.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
As per Claim 3, Paul alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Swaminathan, they collectively teach all of the limitation(s).
Paul and Swaminathan teach Claim 3, which recites: “The method of claim 1, wherein the rendering the selected one or more translated text segments includes: positioning the selected one or more translated text segments locked at a location, in the real-world environment, replacing the at least the portion of the one or more text segments in the first language.”
Paul et al. teaches displaying translated text “at a location corresponding to the original text that has been translated” and “replaces the respective portion of the text ... with the respective translated portion of the text.” Paul et al., ¶ [0547].
For the claimed “locked at a location, in the real-world environment,” Swaminathan et al. teaches an AR tracking algorithm such that “the displayed augmentation information follows (e.g., tracks) the target object,” with pose estimation/tracking maintaining the rendered AR information relative to the target as the camera orientation changes. Swaminathan et al., ¶ [0069]. Thus, the combination teaches replacing the source text with translated text while maintaining the translated content spatially locked to the corresponding real-world target.
The rationale and motivation to combine the references as set forth for claim 1 are incorporated herein by reference for the present claim.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
As per Claim 4, Paul alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Swaminathan, they collectively teach all of the limitation(s).
Paul and Swaminathan teach Claim 4, which recites: “The method of claim 1, wherein the user gaze input tracks movement of at least one eye of a user of the artificial reality system.”
Swaminathan et al. teaches “eye gaze tracking” in which gaze information is based on “the relative position of a user's iris,” with eye and iris coordinates detected and mapped. Swaminathan et al., ¶ [0058]. Swaminathan further teaches creating a tracking effect “as the user's gaze moves across the display.” Swaminathan et al., ¶ [0060]. Thus, Swaminathan teaches tracking eye/iris movement to determine changing user gaze.
The rationale and motivation to combine the references as set forth for claim 1 are incorporated herein by reference for the present claim.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
As per Claim 5, Paul alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Swaminathan, they collectively teach all of the limitation(s).
Paul and Swaminathan teach Claim 5, which recites: “The method of claim 4, wherein the intent to interact is a gaze dwell, within a threshold distance of the at least the portion of the translated one or more text inputs, for greater than a threshold amount of time.”
Swaminathan et al. teaches determining “the amount of time the area of interest 806 lingers on or near an icon 804,” with a “positional tolerance” of approximately “5, 10, or 20 pixels” and linger times of approximately “1, 2, or 3 seconds.” Swaminathan et al., ¶ [0092]. Swaminathan further teaches presenting AR information when an “established linger duration threshold is satisfied.” Swaminathan et al., ¶ [0080]. Thus, Swaminathan teaches a gaze dwell within a threshold distance for greater than a threshold amount of time, applicable to Paul's translated-text indication in the combination.
The rationale and motivation to combine the references as set forth for claim 1 are incorporated herein by reference for the present claim.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
Paul teaches Claim 7, which recites: “The method of claim 1, wherein the rendering the at least the portion of the translated one or more text segments include audibly announcing the at least the portion of the translated one or more text segments.”
Paul et al. teaches that the selected translated word “EGGS” is included in translation card 1470 and that, upon activation of translated-word output control 1470b1, the system “outputs an audible indication (e.g., voice output)” corresponding to the translated word, including “an audible uttering of the translated word.” Paul et al., ¶ [0527]. Thus, Paul directly teaches audibly announcing the selected translated text.
The rationale and motivation to combine the references as set forth for claim 1 are incorporated herein by reference for the present claim.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
Paul teaches Claim 8, which recites: “The method of claim 1, wherein the rendering the at least the portion of the translated one or more text segments includes displaying the at least the portion of the translated one or more text segments.”
Paul et al. teaches that, upon selection of a translated portion, the system “displays ... a first translation user interface object” including “the translation ... of the first portion of the text” without the translation of the second portion. Paul et al., ¶ [0549]. Thus, Paul directly teaches displaying the selected translated text.
The rationale and motivation to combine the references as set forth for claim 1 are incorporated herein by reference for the present claim.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
Paul teaches Claim 9, which recites: “The method of claim 1, further comprising: selecting a rendering mode, for rendering the at least the portion of the translated one or more text segments, based on one or more contextual factors, the one or more contextual factors including one or more preferences for a user of the artificial reality system, one or more situational factors relevant to the real-world environment, or any combination thereof.”
Paul et al. teaches that the “visual appearances” of translated-text objects are “determined by the visual appearance of content (e.g., words, images, background)” of the real-world menu, including “color, texture, size, shape” and font, and may be determined by the “particular underlying content” beneath the translated text. Paul et al., ¶ [0525]. Thus, Paul teaches selecting a rendering mode for translated text based on situational contextual factors of the real-world environment.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
Claim 12 does not include any additional limitations that would significantly distinguish it from claim 1. Therefore, it is likewise rejected under 35 U.S.C. § 103 in view of the same references and for the same reasons set forth above.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
Claim 13 does not include any additional limitations that would significantly distinguish it from claim 4. Therefore, it is likewise rejected under 35 U.S.C. § 103 in view of the same references and for the same reasons set forth above.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
Paul teaches Claim 14, which recites: “The computer-readable storage medium of claim 12, wherein the user input includes gesture input indicative of a gesture made by a user of the artificial reality system, and wherein determining the intent to interact includes determining that the gesture is made within a threshold distance of the at least the portion of the one or more text segments.”
Paul et al. teaches a user-performed “directional gesture” controlling an input representation and determining whether that representation is “within a predetermined distance from a location ... of the detected text.” Paul et al., ¶ [0443]. Thus, Paul teaches determining interaction with detected text based on a gesture-controlled input location being within a threshold distance of the text.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
Claim 15 does not include any additional limitations that would significantly distinguish it from claim 2. Therefore, it is likewise rejected under 35 U.S.C. § 103 in view of the same references and for the same reasons set forth above.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
Claim 16 does not include any additional limitations that would significantly distinguish it from claim 3. Therefore, it is likewise rejected under 35 U.S.C. § 103 in view of the same references and for the same reasons set forth above.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
Claim 17 does not include any additional limitations that would significantly distinguish it from claim 1. Therefore, it is likewise rejected under 35 U.S.C. § 103 in view of the same references and for the same reasons set forth above.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
Claim 18 does not include any additional limitations that would significantly distinguish it from claim 4. Therefore, it is likewise rejected under 35 U.S.C. § 103 in view of the same references and for the same reasons set forth above.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
Claim 19 does not include any additional limitations that would significantly distinguish it from claim 5. Therefore, it is likewise rejected under 35 U.S.C. § 103 in view of the same references and for the same reasons set forth above.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
Ground of Rejection 2
Claim 11 is rejected under 35 U.S.C. § 103 over Paul et al. (US20220337741A1) in view of Swaminathan et al. (US20180253145A1), and further in view of Parkinson et al. (US20130289971A1).
As per Claim 11, Paul alone does not explicitly teach all the limitation(s) of the claim. However, when combined with Swaminathan and Parkinson, they collectively teach all of the limitation(s).
Parkinson teaches Claim 11, which recites: “The method of claim 1, further comprising: determining that the second language is different than the first language, the first language being associated with a user of the artificial reality system; wherein the automatically translating the one or more text segments from the first language to the second language is based on the determination that the second language is different than the first language associated with the user of the artificial reality system.”
Parkinson et al. teaches headsets configured according to the respective user's language, e.g., an “English-speaking user 301A” and a “French-speaking user 301B,” with each headset becoming aware of the other user's language configuration. Parkinson et al., ¶ [0038]. Parkinson further teaches translating a “first language (Language 1)” into a “second language” that is “different from the first/source/one language according to each HCs preferred (or default) language configuration.” Parkinson et al., ¶ [0041]. Thus, Parkinson teaches determining differing user-associated languages and performing translation according to that determination.
Before the effective filing date of the claimed invention, a POSITA would have been motivated to incorporate Parkinson et al.’s user-language determination into Paul et al. and Swaminathan et al. so that text is automatically translated when the detected language differs from the user-associated language, thereby reducing manual input and providing the appropriate translation with predictable results.
PNG
media_image1.png
9
307
media_image1.png
Greyscale
Conclusion
The prior art made of record and relied upon in this action is as follows:
Patent Literature:
Paul et al. (US20220337741A1) — “User interfaces for managing visual content in media.”
Swaminathan et al. (US20180253145A1) — “Enabling augmented reality using eye gaze tracking.”
Parkinson et al. (US20130289971A1) — “Instant Translation System.”
Non-Patent Literature (NPL):
(none)
Note: A PDF copy of each NPL reference is attached with this Office Action. URLs are included for applicant convenience. If a link becomes unavailable in the future, the citation information may be used to locate the reference or access archived versions via the Wayback Machine.
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure and is listed as follows:
Patent Literature:
Yoo (US20150002676A1) — “Smart glass.”
Vukosavljevic et al. (US20150234812A1) — “Text overlay techniques in realtime translation.”
Non-Patent Literature (NPL):
(none)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ADEEL BASHIR whose telephone number is (571) 270-0440. The examiner can normally be reached Monday-Thursday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Hajnik can be reached on (571) 276-7642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ADEEL BASHIR/
Examiner, Art Unit 2616
/DANIEL F HAJNIK/Supervisory Patent Examiner, Art Unit 2616