Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Re claim 1
The limitation of A learning method of a one-shot imitation method based on an model that adaptively one-shot imitates a video demonstration of an expert, the method comprising, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, a one shot imitation method based on a model in the context of this claim encompasses the user forming a mentally imitation model
The limitation of a first stage of learning to infer a semantic skill sequence from a given expert demonstration, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, learning in the context of this claim encompasses the user mentally learning the inference.
The limitation of a second stage of learning to infer dynamics based on state-action pairs, which are the minimum units that form an action trajectory of the expert in the given expert demonstration, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, learning in the context of this claim encompasses the user mentally learning the inference.
The limitation of and a third stage of learning to create an action sequence by combining the inferred semantic skill sequence and dynamics based on given environmental data., as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, learning in the context of this claim encompasses the user mentally learning to create an action sequence.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim only recites one additional element – artificial neural network model. The neural network is recited at a high-level of generality (i.e., as a generic neural network performing a generic function) such that it amounts no more than limiting the claim top the field of neural networks. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using neural network to perform the steps amounts to no more than limiting the claim to a generic neural network. Mere instructions to apply an exception using a generic neural network cannot provide an inventive concept. The claim is not patent eligible.
Re claim 2 claim 2 contains the same abstract idea as claim 1. This judicial exception is not integrated into a practical application. In particular, the claim only recites one additional element – artificial neural network model wherein the artificial neural network model is based on a trained vision-language model configured of a vision encoder and a language encoder. However a artificial neural network model wherein the artificial neural network model is based on a trained vision-language model configured of a vision encoder and a language encoder is conventional see Tschannen US 2024/0169629. This merely limits the claim to use of a well-known neural network (see paragraph 21 “For another example, as described previously, conventional models generally utilize discrete image encoders and text encoders for multimodal vision-language task”). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of a artificial neural network model wherein the artificial neural network model is based on a trained vision-language model configured of a vision encoder and a language encoder is conventional. Mere instructions to apply an exception using a conventional neural network cannot provide an inventive concept. The claim is not patent eligible.
Re claim 3 the limitation of wherein the model is trained based on a prompt, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, training based on a prompt in the context of this claim encompasses the user mentally learning based on a prompt.
Re claim 4 the limitation of wherein, in order to infer parameters about an environment from the state-action pairs, the first stage is trained based on contrastive learning, in which the state-action pairs executed in the same environment are embedded in the same place, and the state-action pairs executed in different environments are embedded far apart from each other., as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, inferring parameters performing contrastive learning and embedding parameters in the context of this claim encompasses the user mentally performing the inference, mentally performing learning based on contrast and mentally embedding parameters in a same and different place.
The analysis with respect to integration into an abstract idea and significantly more is not significantly change from the claim from which this claim depends.
Re claim 5 the limitation of wherein the model comprises: a contrastively trained semantic skill encoder that converts the video demonstration of the expert into the semantic skill sequence; and a semantic skill decoder that infers the optimal skill from the semantic skill sequence depending on a state, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, encoder and decoder in the context of this claim encompasses the user mentally determining a mental encoder and a mentally decoder.
The analysis with respect to integration into an abstract idea and significantly more is not significantly change from the claim from which this claim depends.
Re claim 6 the limitation of wherein the model comprises: model further comprises: skill transfer that infers an action sequence optimized for an operation (execution) of an agent from a given semantic skill sequence and inferred dynamics; and a dynamics encoder that infers the dynamics from expert trajectories, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, skill transfer in the context of this claim encompasses the user mentally performing skill transfer and a user mentally performing dynamics encoding.
The analysis with respect to integration into an abstract idea and significantly more is not significantly change from the claim from which this claim depends.
Re claim 7
The limitation of A one-shot imitation method based on a model that adaptively one-shot imitates a video demonstration of an expert, the method comprising , as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, a one shot imitation method based on a model in the context of this claim encompasses the user forming a mentally imitation model
the limitation of a first stage of inferring a semantic skill sequence from a given expert demonstration, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, inferring in the context of this claim encompasses the user mentally making the inference.
the limitation of a second stage of operating an agent to generate state-action pairs in time series and inferring the dynamics of a current state based thereon, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, learning in the context of this claim encompasses the user mentally generating state action pairs and mentally learning the inference.
the limitation of a third stage of performing an action appropriate for a new domain by combining the inferred semantic skill sequence and inferred dynamics based on currently acquired environmental data and recombining the inferred semantic skill sequence, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, learning in the context of this claim encompasses the user mentally performing the action by combining the skill sequence and the inferred dynamics.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim only recites one additional element – artificial neural network model. The neural network is recited at a high-level of generality (i.e., as a generic neural network performing a generic function) such that it amounts no more than limiting the claim top the field of neural networks. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using neural network to perform the steps amounts to no more than limiting the claim to a generic neural network. Mere instructions to apply an exception using a generic neural network cannot provide an inventive concept. The claim is not patent eligible.
Re claim 8 claim 8 contains the same abstract idea as claim 7. This judicial exception is not integrated into a practical application. In particular, the claim only recites one additional element – artificial neural network model wherein the artificial neural network model is based on a trained vision-language model configured of a vision encoder and a language encoder. However a artificial neural network model wherein the artificial neural network model is based on a trained vision-language model configured of a vision encoder and a language encoder is conventional see Tschannen US 2024/0169629. This merely limits the claim to use of a well-known neural network (see paragraph 21 “For another example, as described previously, conventional models generally utilize discrete image encoders and text encoders for multimodal vision-language task”). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of a artificial neural network model wherein the artificial neural network model is based on a trained vision-language model configured of a vision encoder and a language encoder is conventional. Mere instructions to apply an exception using a conventional neural network cannot provide an inventive concept. The claim is not patent eligible.
Re claim 9 the limitation of wherein the model is trained based on a prompt, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, training based on a prompt in the context of this claim encompasses the user mentally learning based on a prompt.
The analysis with respect to integration into an abstract idea and significantly more is not significantly change from the claim from which this claim depends.
Re claim 10 the limitation of wherein, in order to infer parameters about an environment from the state-action pairs, the first stage is trained based on contrastive learning, in which the state-action pairs executed in the same environment are embedded in the same place, and the state-action pairs executed in different environments are embedded far apart from each other., as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, inferring parameters performing contrastive learning and embedding parameters in the context of this claim encompasses the user mentally performing the inference, mentally performing learning based on contrast and mentally embedding parameters in a same and different place.
The analysis with respect to integration into an abstract idea and significantly more is not significantly change from the claim from which this claim depends.
Re claim 11 the limitation of wherein the model comprises: a contrastively trained semantic skill encoder that converts the video demonstration of the expert into the semantic skill sequence; and a semantic skill decoder that infers the optimal skill from the semantic skill sequence depending on a state, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, encoder and decoder in the context of this claim encompasses the user mentally determining a mental encoder and a mentally decoder.
The analysis with respect to integration into an abstract idea and significantly more is not significantly change from the claim from which this claim depends.
Re claim 12 the limitation of wherein the model comprises: model further comprises: skill transfer that infers an action sequence optimized for an operation (execution) of an agent from a given semantic skill sequence and inferred dynamics; and a dynamics encoder that infers the dynamics from expert trajectories, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, skill transfer in the context of this claim encompasses the user mentally performing skill transfer and a user mentally performing dynamics encoding.
The analysis with respect to integration into an abstract idea and significantly more is not significantly change from the claim from which this claim depends.
Re claim 13
The limitation of model that adaptively one-shot imitates a video demonstration of an expert, wherein in a learning phase, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, a one shot imitation method based on a model in the context of this claim encompasses the user learning a mental imitation model
The limitation of learns to infer a semantic skill sequence from a given expert demonstration, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, learning in the context of this claim encompasses the user mentally learning the inference.
The limitation of learns to infer dynamics based on state-action pairs, which are the minimum units that form an action trajectory of the expert in the given expert demonstration, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, learning in the context of this claim encompasses the user mentally learning the inference.
The limitation of and learns to create an action sequence by combining the inferred semantic skill sequence and dynamics based on given environmental data, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, learning in the context of this claim encompasses the user mentally learning to create an action sequence.
If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea.
This judicial exception is not integrated into a practical application. In particular, the claim only recites additional elements – artificial neural network model A computing device, comprising: a memory that stores an artificial neural network model t; and a processor that executes the artificial neural network model. The processor and memory are recited at a high level of generality such that it amounts to no more than instructions to perform the abstract idea using generic computer components. The neural network is recited at a high-level of generality (i.e., as a generic neural network performing a generic function) such that it amounts no more than limiting the claim top the field of neural networks. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using neural network to perform the steps amounts to no more than limiting the claim to a generic neural network. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using a processor to perform the function amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic neural network combined with generic computer components cannot provide an inventive concept. The claim is not patent eligible.
Re claim 14 The limitation of infers the semantic skill sequence from the given expert demonstration, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, inferring in the context of this claim encompasses the user mentally making the inference.
The limitation of operates an agent to generate the state-action pairs in time series to infer the dynamics of a current state, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, learning in the context of this claim encompasses the user mentally generating state action pairs and mentally learning the inference.
The limitation of a third stage of performs an action appropriate for a new domain by combining the inferred semantic skill sequence and inferred dynamics based on currently acquired environmental data and reassembling the inferred semantic skill sequence, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, learning in the context of this claim encompasses the user mentally performing the action by combining the skill sequence and the inferred dynamics.
The analysis with respect to integration into an abstract idea and significantly more is not significantly change from the claim from which this claim depends.
Re claim 15 claim 15 contains the same abstract idea as claim 13. This judicial exception is not integrated into a practical application. In particular, the claim only recites one additional element – artificial neural network model wherein the artificial neural network model is based on a trained vision-language model configured of a vision encoder and a language encoder. However a artificial neural network model wherein the artificial neural network model is based on a trained vision-language model configured of a vision encoder and a language encoder is conventional see Tschannen US 2024/0169629. This merely limits the claim to use of a well-known neural network (see paragraph 21 “For another example, as described previously, conventional models generally utilize discrete image encoders and text encoders for multimodal vision-language task”). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of a artificial neural network model wherein the artificial neural network model is based on a trained vision-language model configured of a vision encoder and a language encoder is conventional. Mere instructions to apply an exception using a conventional neural network cannot provide an inventive concept. The claim is not patent eligible.
Re claim 16 the limitation of wherein the model is trained based on a prompt, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, training based on a prompt in the context of this claim encompasses the user mentally learning based on a prompt.
The analysis with respect to integration into an abstract idea and significantly more is not significantly change from the claim from which this claim depends.
Re claim 17 the limitation of wherein, in order to infer parameters about an environment from the state-action pairs, the first stage is trained based on contrastive learning, in which the state-action pairs executed in the same environment are embedded in the same place, and the state-action pairs executed in different environments are embedded far apart from each other., as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, inferring parameters performing contrastive learning and embedding parameters in the context of this claim encompasses the user mentally performing the inference, mentally performing learning based on contrast and mentally embedding parameters in a same and different place.
The analysis with respect to integration into an abstract idea and significantly more is not significantly change from the claim from which this claim depends.
Re claim 18 the limitation of wherein the model comprises: a contrastively trained semantic skill encoder that converts the video demonstration of the expert into the semantic skill sequence; and a semantic skill decoder that infers the optimal skill from the semantic skill sequence depending on a state, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, encoder and decoder in the context of this claim encompasses the user mentally determining a mental encoder and a mentally decoder.
The analysis with respect to integration into an abstract idea and significantly more is not significantly change from the claim from which this claim depends.
Re claim 19 the limitation of wherein the model comprises: model further comprises: skill transfer that infers an action sequence optimized for an operation (execution) of an agent from a given semantic skill sequence and inferred dynamics; and a dynamics encoder that infers the dynamics from expert trajectories, as drafted, is a process that, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, skill transfer in the context of this claim encompasses the user mentally performing skill transfer and a user mentally performing dynamics encoding.
The analysis with respect to integration into an abstract idea and significantly more is not significantly change from the claim from which this claim depends.
Re claim 20 Claim 20 contains the same abstract idea as claim 1.
This judicial exception is not integrated into a practical application. In particular, the claim only recites additional elements – artificial neural network model and A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, configure the processor to. The neural network is recited at a high-level of generality (i.e., as a generic neural network performing a generic function) such that it amounts no more than limiting the claim top the field of neural networks. The non-transitory computer-readable storage medium is recited at a high-level of generality (i.e., as a generic medium performing a generic computer function) such that it amounts no more than mere instructions to apply the exception using a generic computer component. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea.
The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using neural network to perform the steps amounts to no more than limiting the claim to a generic neural network. As discussed above with respect to integration of the abstract idea into a practical application, the additional element of using a processor to perform both the ranking and determining steps amounts to no more than mere instructions to apply the exception using a generic computer component. Mere instructions to apply an exception using a generic neural network and generic computer components cannot provide an inventive concept. The claim is not patent eligible.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 1-3, 5-9, and 11-12 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Shin et al “One-shot imitation in a non-stationary environment via multi-modal skill” ICML'23: Proceedings of the 40th International Conference on Machine Learning Article No.: 1307, Pages 31562 – 31578 23 July 2023.
Applicant cannot rely upon the certified copy of the foreign priority application to overcome this rejection because a translation of said application has not been made of record in accordance with 37 CFR 1.55. When an English language translation of a non-English language foreign application is required, the translation must be that of the certified copy (of the foreign application as filed) submitted together with a statement that the translation of the certified copy is accurate. See MPEP §§ 215 and 216.
The examiner notes that the disclosure is made by the join inventors but is more that 1 year prior to the current effective filing date of 8/23/2024.
Re claim 1 Shin discloses
A learning method of a one-shot imitation method based on an artificial neural network model that adaptively one-shot imitates a video demonstration of an expert, the method comprising (see abstract “To tackle the problem, we explore the compositionality of complex tasks, and present a novel skill-based imitation learning framework enabling one-shot imitation and zero-shot adaptation;” note that a one shot learning is used.” And section 1 “n deployment, given a single video demonstration for a new task, the framework translates it to a sequence of se
mantic skills through a sequence encoder built on the vision language model. It then combines the skills with inferred current dynamics to generate actions optimized for not only
the demonstration but also the current dynamics”):
a first stage of learning to infer a semantic skill sequence from a given expert demonstration (see section 3.2 learning semantic skill sequence “A semantic skill sequence encoder Φenc is responsible for mapping an expert demonstration d to a sequence of semantic skills, in that d is segmented into a sequence of dynamics-invariant behavior patterns (action sequences). Each pattern corresponds to a short sequence of expert actions and it can be described in a language instruction in the environment. Such an expert behavior pattern associated with some language instruction is referred to as a semantic skill, as for an expert behavior pattern, its associated language instruction is used to represent expert behaviors on the semantic embedding space of a vision-language aligned model” the first stage corresponds to the training of the skill sequence encoder also see figure 2 and associated caption note that the vision learning model encodes the video into the skill sequence);
a second stage of learning to infer dynamics based on state-action pairs, which are the minimum units that form an action trajectory of the expert in the given expert demonstration (see section 3.1 last paragraph “The contrastive learning procedures are implemented
in the framework, as a semantic skill sequence encoder and a dynamics encoder”
note that second stance corresponds to the dynamics encoder also see figure 2 and associated caption note that the dynamics are inferred from the state action pairs);
and a third stage of learning to create an action sequence by combining the inferred semantic skill sequence and dynamics based on given environmental data (see figure 2 and caption note that the dynamics aware skills are created by combines the skill sequence and the dynamics are used to create a sequence of actions see also section 3.3 “Given a semantic skill sequence encoder Φenc, the skill transfer module πtr is responsible for transferring a semantic skill zt = Φenc(vt:t+H) to an action sequence adapted for environment dynamics”).
Re claim 2 Shin discloses wherein the artificial neural network model is based on a trained vision-language model configured of a vision encoder (see figure 2” In the training phase, a semantic skill set is established from video demonstrations in offline datasets by leveraging a pretrained vision-language model” and section 3.2 note that a semantic skill encoder is uses a vision model) and a language encoder (see section 3.2 “t learning techniques (Zhou et al., 2022). Specifically, we assume that a dataset of expert trajectories D = {τ1,τ2,··· ,τN} contains language instruction data, in that each T-length trajectory τ = {(s1, v1,l1,a1),··· ,(sT,vT,lT,aT)} consists of state s, visual observation v, language instruction l, and action a,
where each language instruction is an element of the instruct tion set L,” note that the dynamics encoder is a language encoder).
Re claim 3 Shin discloses wherein the artificial neural network model is trained based on a prompt. (see section 3.2 “To implement the encoder Φenc, we use the CLIP vision-language pretrained model and sample-efficient prompt learning techniques”).
Re claim 5 Shin discloses wherein the artificial neural network model comprises: a contrastively trained semantic skill encoder that converts the video demonstration of the expert into the semantic skill sequence; and a semantic skill decoder that infers the optimal skill from the semantic skill sequence depending on a state (see figure 3 and caption In (a), the semantic\ skill sequence encoder Φenc and the semantic skill decoder Φdec are trained offline using
The CLIP vision-language pretrained model, where Φenc translates video demonstrations to semantic skill sequences and is contrastively learned, and Φdec learns to infer an optimal skill(from a sequence)upon a state. Note that a contrastively learned encoder encoders into a skill sequence and a decoder infers an optimal skill)
Re claim 6 Shin discloses wherein the artificial neural network model further comprises: skill transfer that infers an action sequence optimized for an operation (execution) of an agent from a given semantic skill sequence and inferred dynamics; and a dynamics encoder that infers the dynamics from expert trajectories (see figure 3 and caption “ In (b), the skill transfer πtr and the dynamics encoder ψenc are trained offline, where πtr learns to infer an action sequence optimized for the deployment setting from a given semantic skill sequence and inferred dynamics ,and ψenc learns to infer dynamics from sub-trajectories”).
Re claim 7 Shin discloses A one-shot imitation method based on an artificial neural network model that adaptively one-shot imitates a video demonstration of an expert, (see abstract “To tackle the problem, we explore the compositionality of complex tasks, and present a novel skill-based imitation learning framework enabling one-shot imitation and zero-shot adaptation;” note that a one shot learning is used.” And section 1 “n deployment, given a single video demonstration for a new task, the framework translates it to a sequence of semantic skills through a sequence encoder built on the vision language model. It then combines the skills with inferred current dynamics to generate actions optimized for not only
the demonstration but also the current dynamics”
the method comprising:
a first stage of inferring a semantic skill sequence from a given expert demonstration (see figure 3 and caption “in(c), for a given demonstration, Φenc first infers a sequence of semantic skills, Φdec infers a current semantic skill” note that skills inferred from the demonstation);
a second stage of operating an agent to generate state-action pairs in time series (see section 3.3 “For H0-length sub-trajectories τt = (st−H0 ,at−H0,...,st−1,at−1), the dynamics encoder
ψenc takes it as input” note that the dynamics encoder uses a times series of stat action pairs as an input) and inferring the dynamics of a current state based thereon (see figure 3 and caption “ψenc infers current dynamics in the non-stationary deployment environment”)
and a third stage of performing an action appropriate for a new domain by combining the inferred semantic skill sequence and inferred dynamics based on currently acquired environmental data and recombining the inferred semantic skill sequence (see introduction paragraph 3 “In deployment, given a single video demonstration for a new task, the framework translates it to a sequence of se mantic skills through a sequence encoder built on the vision
language model. It then combines the skills with inferred current dynamics to generate actions optimized for not only the demonstration but also the current dynamics” note that inferred current dynamics are combined with the semantic skill sequence to generate current dynamics).
Re claim 8 Shin discloses wherein the artificial neural network model is based on a trained vision-language model configured of a vision encoder (see figure 2” In the training phase, a semantic skill set is established from video demonstrations in offline datasets by leveraging a pretrained vision-language model” and section 3.2 note that a semantic skill encoder is uses a vision model ) and a language encoder (see section 3.2 “t learning techniques (Zhou et al., 2022). Specifically, we assume that a dataset of expert trajectories D = {τ1,τ2,··· ,τN} contains language instruction data, in that each T-length trajectory τ = {(s1, v1,l1,a1),··· ,(sT,vT,lT,aT)} consists of state s, visual observation v, language instruction l, and action a,
where each language instruction is an element of the instruct tion set L,” note that the dynamics encoder is a language encoder).
Re claim 9 Shin discloses wherein the artificial neural network model is trained based on a prompt. (see section 3.2 “To implement the encoder Φenc, we use the CLIP vision-language pretrained model and sample-efficient prompt learning techniques”).
Re claim 11 Shin discloses wherein the artificial neural network model comprises: a contrastively trained semantic skill encoder that converts the video demonstration of the expert into the semantic skill sequence; and a semantic skill decoder that infers the optimal skill from the semantic skill sequence depending on a state (see figure 3 and caption In (a), the semantic\ skill sequence encoder Φenc and the semantic skill decoder Φdec are trained offline using
The CLIP vision-language pretrained model, where Φenc translates video demonstrations to semantic skill sequences and is contrastively learned, and Φdec learns to infer an optimal skill (from a sequence)upon a state. Note that a contrastively learned encoder encoders into a skill sequence and a decoder infers an optimal skill)
Re claim 12 Shin discloses wherein the artificial neural network model further comprises: skill transfer that infers an action sequence optimized for an operation (execution) of an agent from a given semantic skill sequence and inferred dynamics; and a dynamics encoder that infers the dynamics from expert trajectories (see figure 3 and caption “ In (b), the skill transfer πtr and the dynamics encoder ψenc are trained offline, where πtr learns to infer an action sequence optimized for the deployment setting from a given semantic skill sequence and inferred dynamics ,and ψenc learns to infer dynamics from sub-trajectories”).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 13-16 and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Shin et al “One-shot imitation in a non-stationary environment via multi-modal skill” ICML'23: Proceedings of the 40th International Conference on Machine Learning Article No.: 1307, Pages 31562 – 31578 23 July 2023 in view of Cherian US 20240069501 A1.
Re claim 13 Shin discloses 13. an artificial neural network model that adaptively one-shot imitates a video demonstration of an expert; (see abstract “To tackle the problem, we explore the compositionality of complex tasks, and present a novel skill-based imitation learning framework enabling one-shot imitation and zero-shot adaptation;” note that a one shot learning is used.” And section 1 “n deployment, given a single video demonstration for a new task, the framework translates it to a sequence of semantic skills through a sequence encoder built on the vision language model. It then combines the skills with inferred current dynamics to generate actions optimized for not only the demonstration but also the current dynamics”), wherein in a learning phase, the artificial neural network model:
learns to infer a semantic skill sequence from a given expert demonstration; (see section 3.2 learning semantic skill sequence “A semantic skill sequence encoder Φenc is responsible for mapping an expert demonstration d to a sequence of semantic skills, in that d is segmented into a sequence of dynamics-invariant behavior patterns (action sequences). Each pattern corresponds to a short sequence of expert actions and it can be described in a language instruction in the environment. Such an expert behavior pattern associated with some language instruction is referred to as a semantic skill, as for an expert behavior pattern, its associated language instruction is used to represent expert behaviors on the semantic embedding space of a vision-language aligned model” the first stage corresponds to the training of the skill sequence encoder also see figure 2 and associated caption note that the vision learning model encodes the video into the skill sequence);
learns to infer dynamics based on state-action pairs, which are the minimum units that form an action trajectory of the expert in the given expert demonstration; (see section 3.1 last paragraph “The contrastive learning procedures are implemented
in the framework, as a semantic skill sequence encoder and a dynamics encoder”
note that second stance corresponds to the dynamics encoder also see figure 2 and associated caption note that the dynamics are inferred from the state action pairs);
and learns to create an action sequence by combining the inferred semantic skill sequence and dynamics based on given environmental data. (see figure 2 and caption note that the dynamics aware skills are created by combines the skill sequence and the dynamics are used to create a sequence of actions see also section 3.3 “Given a semantic skill sequence encoder Φenc, the skill transfer module πtr is responsible for transferring a semantic skill zt = Φenc(vt:t+H) to an action sequence adapted for environment dynamics”).
Shin does not expressly disclose
A computing device, comprising: a memory that stores an artificial neural network model, and a processor that executes the artificial neural network model.
In a similar field of endeavor, Cherian discloses A computing device, comprising: a memory that stores an artificial neural network model, (see paragraph 77 “he memory 206 may include random access memory (RAM), read only memory (ROM), flash memory, or any other suitable memory systems. The processor 205 is connected through the bus 201 to one or more input interfaces and the other devices. In an embodiment, the memory 206 is embodied within the controller 209 and may additionally store the hierarchical multimodal RL neural network 209a” note that the memory may store the neural network model) and a processor that executes the artificial neural network model (see paragraph 76 note that the processor executes the neural network model.) One of ordinary skill in the art could have easily implemented the method of Shin using the processor and memory of Cherian the results would be the same as the computer would merely be used as a tool to implement the invention. Further such as processor and memory are merely used as a tool to implement the process of Shin the elements do the same thing separately as the do in combination. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Shin and Cherian.
Re claim 14 Shin further discloses wherein, in an execution phase, the artificial neural network model: infers the semantic skill sequence from the given expert demonstration (see figure 3 and caption “in(c), for a given demonstration, Φenc first infers a sequence of semantic skills, Φdec infers a current semantic skill);
operates to generate state-action pairs in time series (see section 3.3 “For H0-length sub-trajectories τt = (st−H0, at−H0,...,st−1,at−1), the dynamics encoder ψenc takes it as input” note that the dynamics encoder uses a times series of stat action pairs as an input) and inferring the dynamics of a current state based thereon (see figure 3 and caption “ ψenc infers current dynamics in the non-stationary deployment environment”)
; and
performs an action appropriate for a new domain by combining the inferred semantic skill sequence and inferred dynamics based on currently acquired environmental data and reassembling the inferred semantic skill sequence (see introduction paragraph 3 “In deployment, given a single video demonstration for a new task, the framework translates it to a sequence of se mantic skills through a sequence encoder built on the vision
language model. It then combines the skills with inferred current dynamics to generate actions optimized for not only the demonstration but also the current dynamics” note that inferred current dynamics are combined with the semantic skill sequence to generate current dynamics).
Re claim 15 Shin discloses wherein the artificial neural network model is based on a trained vision-language model configured of a vision encoder (see figure 2” In the training phase, a semantic skill set is established from video demonstrations in offline datasets by leveraging a pretrained vision-language model” and section 3.2 note that a semantic skill encoder is uses a vision model) and a language encoder (see section 3.2 “. Specifically, we assume that a dataset of expert trajectories D = {τ1,τ2,··· ,τN} contains language instruction data, in that each T-length trajectory τ = {(s1, v1,l1,a1),··· ,(sT,vT,lT,aT)} consists of state s, visual observation v, language instruction l, and action a, where each language instruction is an element of the instruct tion set L,” note that the dynamics encoder is a language encoder).
Re claim 16 Shin discloses wherein the artificial neural network model is trained based on a prompt. (see section 3.2 “To implement the encoder Φenc, we use the CLIP vision-language pretrained model and sample-efficient prompt learning techniques”).
Re claim 18 Shin discloses wherein the artificial neural network model comprises: a contrastively trained semantic skill encoder that converts the video demonstration of the expert into the semantic skill sequence; and a semantic skill decoder that infers the optimal skill from the semantic skill sequence depending on a state (see figure 3 and caption In (a), the semantic\ skill sequence encoder Φenc and the semantic skill decoder Φdec are trained offline using
The CLIP vision-language pretrained model, where Φenc translates video demonstrations to semantic skill sequences and is contrastively learned, and Φdec learns to infer an optimal skill(from a sequence)upon a state. Note that a contrastively learned encoder encoders into a skill sequence and a decoder infers an optimal skill)
Re claim 19 Shin discloses wherein the artificial neural network model further comprises: skill transfer that infers an action sequence optimized for an operation (execution) of an agent from a given semantic skill sequence and inferred dynamics; and a dynamics encoder that infers the dynamics from expert trajectories (see figure 3 and caption “ In (b), the skill transfer πtr and the dynamics encoder ψenc are trained offline, where πtr learns to infer an action sequence optimized for the deployment setting from a given semantic skill sequence and inferred dynamics ,and ψenc learns to infer dynamics from sub-trajectories”).
Re claim 20 Shin disclose the method of claim 1. Shin does not expressly disclose A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, configure the processor to perform. Cherian discloses A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, configure the processor to perform (see paragraph 76 “The entity 150 includes the processor 205 configured to execute stored instructions, as well as a memory 206 that stores instructions that are executable by the processor 205. The processor 205 may be a single core processor, a multi-core processor, a computing cluster, or any number of other configurations.” Note that the memory corresponds to the medium which is executed by the processor). One of ordinary skill in the art could have easily implemented the method of Shin using the processor and memory of Cherian the results would be the same as the computer would merely be used as a tool to implement the invention. Further such as processor and memory are merely used as a tool to implement the process of Shin the elements do the same thing separately as they do in combination. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine Shin and Cherian.
Allowable Subject Matter
Claim 4 10 and 17 rejected under 35 U.S.C. 101, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims and the rejection under 35 U.S.C. 101 were overcome.
Re claim 4 Shin discloses wherein, in order to infer parameters about an environment from the state-action pairs, the first stage is trained based on contrastive learning (see section 3.1 last paragraph “The contrastive learning procedures are implemented in the framework, as a semantic skill sequence encoder and a dynamics encoder, and the meta-learning procedure is
implemented as a skill transfer module. We explain these modules in Figure 3, and Sections 3.2 and 3.3” see also figure 2 and 3 note that the dynamics encoder takes in states and actions and outputs dynamics), however does not expressly disclose in which the state-action pairs executed in the same environment are embedded in the same place, and the state-action pairs executed in different environments are embedded far apart from each other.
Claims 10 and 17 contain similar subject matter.
Cited Art
The following is a recitation of the art considered relevant by not used in a rejection above
Pertsch er al “Cross-Domain Transfer via Semantic Skill Imitation” ArXiv 2022 discloses We propose an approach for semantic imitation, which uses demonstrations from a source domain, e.g., human videos, to accelerate reinforcement
learning (RL) in a different target domain, e.g., a robotic manipulator in a simulated
kitchen. Instead of imitating low-level actions like joint velocities, our approach
imitates the sequence of demonstrated semantic skills like “opening the microwave”
or “turning on the stove”. This allows us to transfer demonstrations across environments (e.g., real-world to simulated kitchen) and agent embodiments (e.g., bimanual human demonstration to robotic arm). We evaluate on three challenging cross-domain learning problems and match the performance of demonstration accelerated RL approaches that require in-domain demonstrations. In a simulated kitchen environment, our approach learns long-horizon robot manipulation tasks, using less than 3 minutes of human video demonstrations from a real-world kitchen. This enables scaling robot learning via the reuse of demonstrations, e.g., collected
as human videos, for learning in any number of target domains.”
Hakhamaneshi et al “HIERARCHICAL FEW-SHOT IMITATION WITH SKILL TRANSITION MODELS” arXiv:2107.08981 March 10 2022 discloses
A desirable property of autonomous agents is the ability to both solve long-horizon problems and generalize to unseen tasks. Recent advances in data-driven skill learning have shown that extracting behavioral priors from offline data can enable agents to solve challenging long-horizon tasks with reinforcement learning. However, generalization to tasks unseen during behavioral prior training remains an outstanding challenge. To this end, we present Few-shot Imitation with Skill Transition Models (FIST), an algorithm that extracts skills from offline data and utilizes them to generalize to unseen tasks given a few downstream demonstrations. FIST learns an inverse skill dynamics model, a distance function, and utilizes a semi-parametric approach for imitation. We show that FIST is capable of generalizing to new tasks and substantially outperforms prior baselines in navigation experiments requiring traversing unseen parts of a large maze and 7-DoF robotic arm experiments requiring manipulating previously unseen objects in a kitchen.
Jiang et al VIMA: General Robot Manipulation with Multimodal Prompts arXiv:2210.03094 May 28 2023 discloses
Prompt-based learning has emerged as a successful paradigm in natural language processing, where a single general-purpose language model can be instructed to perform any task specified by input prompts. Yet task specification in robotics comes in various forms, such as imitating one-shot demonstrations, following language instructions, and reaching visual goals. They are often considered different tasks and tackled by specialized models. We show that a wide spectrum of robot manipulation tasks can be expressed with multimodal prompts, interleaving textual and visual tokens. Accordingly, we develop a new simulation benchmark that consists of thousands of procedurally-generated tabletop tasks with multimodal prompts, 600K+ expert trajectories for imitation learning, and a four-level evaluation protocol for systematic generalization. We design a transformer-based robot agent, VIMA, that processes these prompts and outputs motor actions autoregressively. VIMA features a recipe that achieves strong model scalability and data efficiency. It outperforms alternative designs in the hardest zero-shot generalization setting by up to 2.9× task success rate given the same training data. With 10× less training data, VIMA still performs 2.7× better than the best competing variant.
(See abstract.)
Hori US 20240288870 A1 discloses A method, a system and a computer program product are provided for applying a neural network including an action sequence decoder for generating an action sequence for a robot to perform a task. The neural network is applied to generate the action sequence based on recordings demonstrating humans performing tasks. In an example, the method comprises collecting a recording and a sequence of captions describing scenes in the recording; extracting feature data from the recording; encoding the extracted feature data to produce a sequence of encoded features; and applying the action sequence decoder to produce a sequence of actions for the robot based on the sequence of encoded features having a semantic meaning corresponding to a semantic meaning of the sequence of captions. The feature data includes features of a video signal, an audio signal, and/or text transcription capturing a performance of the task. (see abstract)
Hasenclever US 20200104685 discloses A computer-implemented method of training a student machine learning system comprises receiving data indicating execution of an expert, determining one or more actions performed by the expert during the execution and a corresponding state-action Jacobian, and training the student machine learning system using a linear-feedback-stabilized policy. The linear-feedback-stabilized policy may be based on the state-action Jacobian. Also a neural network system for representing a space of probabilistic motor primitives, implemented by one or more computers. The neural network system comprises an encoder configured to generate latent variables based on a plurality of inputs, each input comprising a plurality of frames, and a decoder configured to generate an action based on one or more of the latent variables and a state. (see abstract).
Chaudhury US 20190385061 A1 discloses “A computer-implemented method is provided for learning an action policy. The method includes obtaining, by a processor, environment dynamics including triplets of a state, an action, and a next state. The state in each of the triplets is an expert state. The method further includes training, by the processor using the environment dynamics as training data, a dynamics model which obtains a pair of the state and the action as an input and outputs, for each next state, state-transition probabilities. The method also includes learning, by the processor, the action policy using trajectories of expert states according to a supervised learning technique by back-propagating error gradients through the trained dynamics model.” (see abstract).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SEAN T MOTSINGER whose telephone number is (571)270-1237. The examiner can normally be reached 9AM-5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at (571) 272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SEAN T MOTSINGER/Primary Examiner, Art Unit 2673