Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 7-8 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al (US20260245300) in view of Perel et al (US20260087757).
Regarding Claim 1. Liu teaches A method comprising:
inputting, by a processing device, an object description into a large language model that creates an initial compact graph representing an initial hierarchy of initial object attributes (Liu, abstract, the invention describes methods to segment, using at least one machine learning model, a first representation of an object into a plurality of parts according to a hierarchy, where the hierarchy indicates groups of datapoints for at least one part of the plurality of parts based on a value. The one or more processors generate at least one second representation of the object based on the at least one of the part of the plurality of parts.
[0032] Approaches in accordance with various embodiments can be used to generate one or more parameters for a content generation environment. In at least one embodiment, a trained machine learning (ML) and/or artificial intelligence (AI) system, such as a large language model (LLM) or a vision language model (VLM), may be used to generate parameters for the content generation environment, such as, but not limited to, camera settings, scene lighting, video parameters, and/or the like, used for displaying objects within a scene. The parameters may be based on an input provided by a user or a proxy for a user to a trained language model (e.g., LLM, VLM, etc.) that can then generate one or more settings in accordance with the input.);
Liu fails to explicitly teach, however, Perel teaches based on the object description (Perel, abstract, the invention describes methods for reconstructing, segmenting, and/or simulating pipeline. A first computing system can obtain at least one object segmented from video data. The first computing system can densify the at least one object. The first computing system can simulate one or more interactions of the voxelized volume of the at least one densified object to update at least one physical attribute of the at least one object. The first computing system can generate at least one image depicting at least a portion of the at least one object with the at least one updated physical attribute.
[0152] In at least some implementations, language models, such as large language models (LLMs), vision language models (VLMs), multi-modal language models (MMLMs), and/or other types of generative artificial intelligence (AI) can be implemented. Generally, the language models can be used to process, analyze, and generate multi-modal content (e.g., text, images, video, 3D models) in various applications, such as those within the 3D RCS pipeline described above. … For example, vision language models (VLMs), or more generally multi-modal language models (MMLMs ), can be implemented to accept image, video, audio, textual, 3D design (e.g., CAD), and/or other inputs data types and/or to generate or output image, video, audio, textual, 3D design, and/or other output data types.).
Liu and Perel are analogous art because they both teach method of AI generating image objects. Liu further teaches generating a hierarchical structure for multiple parts of the image object. Perel further teaches using natural language model to read the description of the object and generating corresponding output. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AI method of generating image object (taught in Liu), to further use natural language description as input to extract features from the text (taught in Perel), so as to provide detailed and realistic 3D reconstructions (Perel, [0001-0002]).
The combination of Liu and Perel further teaches generating, by the processing device, an initial object model based on the initial compact graph (Liu, [0079] FIG. 9B is a flow diagram showing a method 910 for segmenting an object (e.g., 3D volume 300) into a plurality of parts (e.g., parts 136) and generating a representation of the object based on at least one part of the plurality of parts, in accordance with some implementations of the present disclosure. At block 912, the method 910 include segmenting a first representation of an object into a plurality of parts according to a hierarchy (e.g., hierarchical K-means clustering). The hierarchy can indicate groups of datapoints (e.g., points) for at least one part of the plurality of parts based on a value. The value can be a granularity value and can indicate a number of parts for which to segment the first representation of the object into.);
inputting, by the processing device, an object edit into the large language model that creates an updated compact graph representing an updated hierarchy of updated object attributes (Liu, [0029], … The simulation environment may be built into applications or software platforms for creating, generating, modifying, or manipulating 3D models or parts.
[0052] As an example, responsive to the clustering algorithm being agglomerative clustering, each data point (e.g., point) on the part feature field 132 is treated as an individual cluster. The distance between each pair of clusters is then calculated. A computation of all the distances of each pair of clusters in the part feature field 132 may be represented by a distance matrix. The distance may be calculated by, for example, Euclidean or Manhattan distance. Based on the distance, the two clusters that are closest to each other are merged together to reduce a number of clusters of the part feature field 132. After merging the two clusters, the distance matrix can be updated to reflect distances between the merged cluster and remaining clusters. To compute the distances between these clusters, the feature field part generator 134 may use at least one of single linkage, complete linkage, average linkage, or Ward's method. Merging two clusters and updating the distance matrix is then repeated until either all data points on the part feature field 132 are in a single cluster, or until a threshold is met. The threshold can be a predefined number of clusters (e.g., predefined by the user). A hierarchical tree (e.g., dendrogram) can also be formed based on the clustering process. The hierarchical tree can be cut at a specified level to obtain a desired number of clusters.
[0081] At block 914, the method 910 includes generating at least one second representation of the object based on at least one part of the plurality of parts. The first representation of the object and the at least one second representation of the object can be 3D. In various implementations, the method 910 can include adjusting the at least one second representation according to the at least one part of the plurality of parts. For example, the at least one second representation can be edited following generation using the at least one of the plurality of parts.
Therefore, the hierarchical tree of parts of the object can be changed by updating the distance matrix and/or clustering.); and
replacing, by the processing device, the initial object model with an updated object model generated based on the updated compact graph (Liu, [0052] As an example, responsive to the clustering algorithm being agglomerative clustering, each data point (e.g., point) on the part feature field 132 is treated as an individual cluster. The distance between each pair of clusters is then calculated. A computation of all the distances of each pair of clusters in the part feature field 132 may be represented by a distance matrix. The distance may be calculated by, for example, Euclidean or Manhattan distance. Based on the distance, the two clusters that are closest to each other are merged together to reduce a number of clusters of the part feature field 132. After merging the two clusters, the distance matrix can be updated to reflect distances between the merged cluster and remaining clusters. To compute the distances between these clusters, the feature field part generator 134 may use at least one of single linkage, complete linkage, average linkage, or Ward's method. Merging two clusters and updating the distance matrix is then repeated until either all data points on the part feature field 132 are in a single cluster, or until a threshold is met. The threshold can be a predefined number of clusters (e.g., predefined by the user). A hierarchical tree (e.g., dendrogram) can also be formed based on the clustering process. The hierarchical tree can be cut at a specified level to obtain a desired number of clusters.
Therefore, the hierarchy tree structure is updated when distance matrix and/or clustering are changed.).
Regarding Claim 2. The combination of Liu and Perel further teaches The method of claim 1, wherein the updated compact graph is created by modifying the initial hierarchy or the initial object attributes in response to inputting the object edit and the initial compact graph into the large language model (Liu, [0051] The feature field part generator 134 may segment the part feature field 132 into at least two parts 136. In some implementations, a user can determine a granularity of the feature field part generator 134. For example, the user can determine how many parts 136 the feature field part generator 134 segments the part feature field 132 into. As a result, the threshold of the clustering algorithm may change based on the user selection of granularity. The feature field part generator 134 can thus segment the part feature field 132 hierarchically based on the granularity and a desired number of the parts 136. The hierarchy can indicate which of the feature vectors to group together according to the granularity value. The granularity value can indicate, for example, a threshold for a difference between at least one of a similarity or distance between each of the feature vectors. The feature vectors that are at or below the threshold can be grouped (e.g., clustered) into one cluster, indicating a part 136. For example, the threshold may be lower for a desired number of parts 136 being 2 (e.g., lower granularity) and higher for a desired number of parts 136 being 8 (e.g., higher granularity).).
Regarding Claim 7. The combination of Liu and Perel further teaches The method of claim 1, wherein:
the generating includes interpreting, by the processing device, the initial compact graph by forming an initial set of geometric primitives that are combined into the initial object model; and
the replacing includes interpreting, by the processing device, the updated compact graph by forming an updated set of geometric primitives that are combined into the updated object model (Liu, [0052] As an example, responsive to the clustering algorithm being agglomerative clustering, each data point (e.g., point) on the part feature field 132 is treated as an individual cluster. The distance between each pair of clusters is then calculated. A computation of all the distances of each pair of clusters in the part feature field 132 may be represented by a distance matrix. The distance may be calculated by, for example, Euclidean or Manhattan distance. Based on the distance, the two clusters that are closest to each other are merged together to reduce a number of clusters of the part feature field 132. After merging the two clusters, the distance matrix can be updated to reflect distances between the merged cluster and remaining clusters. To compute the distances between these clusters, the feature field part generator 134 may use at least one of single linkage, complete linkage, average linkage, or Ward's method. Merging two clusters and updating the distance matrix is then repeated until either all data points on the part feature field 132 are in a single cluster, or until a threshold is met. The threshold can be a predefined number of clusters (e.g., predefined by the user). A hierarchical tree (e.g., dendrogram) can also be formed based on the clustering process. The hierarchical tree can be cut at a specified level to obtain a desired number of clusters.
[0056] The system 200 may include the 3D representation generator 102, the feature space generator 110, the combiner 122, the dimensionality expander 126, and the feature field generator 130 as described with respect to the system 100. The 3D representation generator 102 may receive an input training geometry 202 indicative of a training 3D volume. The training 3D volume may be included in a training dataset. The training dataset may include 3D labels and 2D labels of a plurality of training 3D volumes. The 3D labels may be derived from an annotated shape dataset without textures with information on internal structures of the training 3D volumes. The 3D labels may be derived from an annotated shape dataset without textures with information on internal structures of the training 3D volumes. The 2D labels may be derived from multi-view rendering of the training 3D volumes by a 2D segmentation model. Prior to updating (e.g., training), the feature field generator 130 may include a plurality of priors and a plurality of weights which may be randomly selected. The plurality of priors can include 3D data, data from 2D foundation models, and heuristics such as convexity, geometric primitives, or color, among others. The plurality of weights may be iteratively updated during the updating process until the feature field generator 130 reaches convergence.
Therefore, the hierarchy tree structure is updated when distance matrix and/or clustering are changed. Such updating including change regrouping of the different parts/shapes/primitives into updated 3D model.).
Regarding Claim 8. The combination of Liu and Perel further teaches The method of claim 1, wherein the initial object model and the updated object model comprise three dimensional mesh representations of an object (Liu, [0029] The systems and methods in accordance with the present disclosure may be applied to various technical fields, such as 3D shape generation, simulation, autonomous driving, animation, and editing AI-generated meshes. For example, systems and methods in accordance with the present disclosure enable generation of meshes per part instead of an entire object. In some implementations, the systems and methods described herein may provide segmentation data to a simulation environment (e.g., NVIDIA's DriveSIM) for the simulation environment to identify regions of interest (e.g., legs, arms) to perform operations associated with the segments of the 3D part. The simulation environment may be built into applications or software platforms for creating, generating, modifying, or manipulating 3D models or parts.
[0037] Once the 3D representation generator 102 receives the input geometry 104, the 3D representation generator 102 can generate at least one 3D representation. The 3D representation can model and represent the geometry of the 3D volume. The 3D representation can have different formats such as, but not limited to, a signed distance function (SDF) volume (e.g., encoded distance for each point of the input geometry 104), a point cloud (e.g., collection of spatial coordinates of the input geometry 104), a mesh representation (e.g., including vertices, edges, faces), a volumetric pixel (e.g., voxel) grid, or a parameter point model.).
Claims 3, 5-6, 9 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al (US20260245300) in view of Perel et al (US20260087757) further in view of Achlioptas et al. ("ShapeTalk: A language dataset and framework for 3d shape edits and deformations." 2023 IEEE).
Regarding Claim 3. The combination of Liu and Perel fails to explicitly teach, however, Achlioptas teaches The method of claim 1, further comprising:
displaying, by the processing device, a user interface at a display device that includes an object preview window including an initial rendered image that depicts a view of the initial object model; and
responsive to receiving the object edit via the user interface, modifying, by the processing device, the object preview window by displaying a subsequent rendered image that depicts an updated view of the updated object model (Achlioptas, abstract, the paper describes the most extensive existing corpus of natural language utterances describing shape differences: ShapeTalk. ShapeTalk contains over half a million discriminative utterances produced by contrasting the shapes of common 3D objects for a variety of object classes and degrees of similarity. We also introduce a generic framework, ChangeIt3D, which builds on ShapeTalk and can use an arbitrary 3D generative model of shapes to produce edits that align the output better with the edit or deformation description. Finally, we introduce metrics for the quantitative evaluation of language-assisted shape editing methods that reflect key desiderata within this editing setup. We note that ShapeTalk allows methods to be trained with explicit 3D-to-language data, bypassing the necessity of “lifting” 2D to 3D using methods like neural rendering, as required by extant 2D image-language foundation models.
Page 8, col 1, par 2, … Figure 4 shows qualitative examples of decoded shape-edits with our ImNet-based models. These examples showcase the ability of ShapeTalk and ChangeIt3D to give rise to nuanced yet semantically correct edits to shapes. Moreover, our method appears to be able to preserve the overall identity of the input shape, and oftentimes create localized/minimal shape edits that can cover both structural and continuous changes.).
Liu, Perel and Achlioptas are analogous art because they all teach method of AI generating image objects. Liu further teaches generating a hierarchical structure for multiple parts of the image object. Perel further teaches using natural language model to read the description of the object and generating corresponding output. Achlioptas further teaches GUI to preview the 3D object model before and after editing. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AI method of generating image object (taught in Liu and Perel), to further use the GUI provided to the user to further update the 3D object model using natural language (taught in Achlioptas), so as to provide convenient tool to help visually-impaired users to interact with object of interest and change them to better fit their design needs (Achlioptas, page 1, col 2, par 3).
Regarding Claim 5. The combination of Liu, Perel and Achlioptas further teaches The method of claim 1, further comprising:
generating, by the processing device, renderable data that captures a view of the initial object model or the updated object model (Liu, [0029] The systems and methods in accordance with the present disclosure may be applied to various technical fields, such as 3D shape generation, simulation, autonomous driving, animation, and editing AI-generated meshes. For example, systems and methods in accordance with the present disclosure enable generation of meshes per part instead of an entire object. In some implementations, the systems and methods described herein may provide segmentation data to a simulation environment (e.g., NVIDIA's DriveSIM) for the simulation environment to identify regions of interest (e.g., legs, arms) to perform operations associated with the segments of the 3D part. The simulation environment may be built into applications or software platforms for creating, generating, modifying, or manipulating 3D models or parts.); and
outputting, by the processing device, a rendered image based on the renderable data for display at a display device (Liu, [0037] Once the 3D representation generator 102 receives the input geometry 104, the 3D representation generator 102 can generate at least one 3D representation. The 3D representation can model and represent the geometry of the 3D volume. The 3D representation can have different formats such as, but not limited to, a signed distance function (SDF) volume (e.g., encoded distance for each point of the input geometry 104), a point cloud (e.g., collection of spatial coordinates of the input geometry 104), a mesh representation (e.g., including vertices, edges, faces), a volumetric pixel (e.g., voxel) grid, or a parameter point model.).
Regarding Claim 6. The combination of Liu, Perel and Achlioptas further teaches The method of claim 1, wherein the object edit comprises a natural language instruction to modify one or more visual characteristics of the initial object model device (Achlioptas, As shown in Figure 2, user’s language “The backrest is comprised of two flat rectangular panels, separated by a thin space for a modern design” is analyzed by language model and change initial 3D model chair on the left to a modern look chair on the right.).
The reasoning for combination of Liu, Perel and Achlioptas is the same as described in Claim 3.
Regarding Claim 9. The combination of Liu, Perel and Achlioptas further teaches The method of claim 1, wherein the initial object attributes and the updated object attributes comprise geometric properties, material properties, or both geometric and material properties of the initial object model and the updated object model, respectively (Achlioptas, page 5, col 2, par 2, For modularity, we train the editing process via a two stage approach. As the generative model G needs to capture sufficient geometric information, we pretrain an autoencoder during the first stage to achieve good reconstructions. Once G is pretrained, L is trained to associate higher utterance-compatibility with the target shape than with the distractor shape, using the latent representations given by the pretrained network G as input.
Perel, [0095] In some implementations, the simulation stage can refer to the stage in the 3D RCS pipeline in which the system 100 simulates interactions of the voxelized volume. That is, the simulator 120 can inject a plurality of volumetric elements (e.g., isotropic Gaussians) in the interior of the voxelized volume to populate the interior. The simulator 120 can simulate one or more interactions of the voxelized volume of the densified object to update at least one physical attribute, such as rigidity or elasticity, of the object. The system 100 can include at least one simulator 120. The simulator 120 can include any one or more physics-based models, rules, heuristics, algorithms, functions, or various combinations thereof to perform operations including simulating one or more physical interactions (e.g., rigid body dynamics, elasticity) of the object.).
The reasoning for combination of Liu, Perel and Achlioptas is the same as described in Claim 3.
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Liu et al (US20260245300) in view of Perel et al (US20260087757), Achlioptas et al. further in view of Hizmi (“LLMto3D Generating parametric Objects using Large Language Model”, 2024, ACADIA).
Regarding Claim 4. The combination of Liu, Perel and Achlioptas fails to explicitly teach, however, Hizmi teaches The method of claim 3, further comprising:
receiving, by the processing device, the object edit by displaying a set of parameter controls indicating adjustable attributes of the initial object model and interpreting the object edit based on user inputs received at the set of parameter controls (Hizmi, abstract, the paper describes methods of using Machine Learning (ML) to generate 3D objects from textual descriptions. We introduce a novel method for translating natural language descrip-tions into parametric 3D objects using Large Language Models (LLMs). Our approach employs multiple agents, each one an LLM pre-trained for a specific task. The first agent deconstructs textual prompts into design elements and describes their geometry and spatial relations. The second agent translates the description into code using the Rhino. Geometry coding library in the Rhino3D-Grasshopper modeling environment. A final agent reassembles the models and adds parametric control interfaces, enabling customizable outputs. In this paper, we describe the method's architecture, and the training methodol-ogies used to fine-tune the models.
Page 7, col 2, par 2, Code Output: The output of the codes in the example dataset is a Python code describing a BREP geometry, which is developed by the LLM agents. This code is then transmitted back to the Grasshopper environment which realizes the geometry via the Hops components. In addition to the geometry, the process automatically generates input sliders through an interface invoked by the Python scripts and executed by the Hops components. These sliders enable users to tailor the final object precisely. An illustrative example demonstrates a kettle produced by this method (Figure 8). The output parameters allow for extensive manipulation to achieve the desired design and functionality of the kettle.).
Liu, Perel, Achlioptas and Hizmi are analogous art because they all teach method of AI generating image objects. The combination of Liu Perel and Achlioptas further teaches generating a hierarchical structure for multiple parts of the image object using LLM/VLM. Hizmi further teaches parameter control interface for user to edit the 3D object model. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AI method of generating image object (taught in Liu, Perel, Achlioptas), to further use the parameter control interface (taught in Hizmi), so as to provide intuitive user interface for use to further customize 3D object for printability or manufacturability (Hizmi, abstract).
Claims 10-11 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al (US20260245300) in view of Chen et al (CN117253237), Achlioptas et al. ("ShapeTalk: A language dataset and framework for 3d shape edits and deformations." 2023 IEEE).
Regarding Claim 10. Liu teaches A method comprising:
providing, by a processing device, a dataset comprising text descriptions of objects and corresponding compact graphs that represent hierarchies of object attributes (Liu, abstract, the invention describes methods to segment, using at least one machine learning model, a first representation of an object into a plurality of parts according to a hierarchy, where the hierarchy indicates groups of datapoints for at least one part of the plurality of parts based on a value. The one or more processors generate at least one second representation of the object based on the at least one of the part of the plurality of parts.
[0032] Approaches in accordance with various embodiments can be used to generate one or more parameters for a content generation environment. In at least one embodiment, a trained machine learning (ML) and/or artificial intelligence (AI) system, such as a large language model (LLM) or a vision language model (VLM), may be used to generate parameters for the content generation environment, such as, but not limited to, camera settings, scene lighting, video parameters, and/or the like, used for displaying objects within a scene. The parameters may be based on an input provided by a user or a proxy for a user to a trained language model (e.g., LLM, VLM, etc.) that can then generate one or more settings in accordance with the input.).
Liu fails to explicitly teach, however, Chen teaches a dataset comprising text descriptions of objects (Chen, abstract, the invention describes a method comprising the steps of obtaining a to-be-classified image, and determining a multidimensional text associated with the to-be-classified image; based on the texts of the multiple dimensions, determining text feature representations of the multiple dimensions; according to the multidimensional text feature representation, an extended text of the to-be-classified image is generated, and semantic information expressed by the extended text is more than semantic information expressed by the multi-dimensional text; extracting text semantic features from the extended text to obtain extended text semantic features; extracting image semantic features of the to-be-classified images; and based on the extended text semantic feature and the image semantic feature,
classifying the to-be-classified image to obtain a category to which the to-be-classified image belongs. By adopting the method, the accuracy of image classification can be improved.
[0106] When multiple dimensions include text recognition, character detection is performed on the image to be classified to detect the character regions in the image and determine the location information of the character regions. Based on the location information of the character regions, each character region is segmented from the image to be classified, and character recognition is performed on each segmented character region to obtain the recognition result of each character region. The recognition result of each character region is used as the recognized text of the image to be classified in the text recognition dimension. In this embodiment, the location information is, for example, the coordinates of a character region.
Therefore, the text description is used to detect object segmentation. Further see Figure 3.);
Liu and Chen are analogous art because they both teach method of using natural language (text description) to identify part of interest in a scene. Liu further teaches generating a hierarchical structure for multiple parts of the image object. Chen further teaches using Large Language Model to summarize text description from different dimensions and detect object segmentation of interest. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AI method of generating image object (taught in Liu), to further use LLM to summarize text description of multi-dimensions (taught in Chen), so as to provide method to accurately classify image objects (Chen, [0001-0003]).
The combination of Liu and Chen fails to explicitly teach, however, Achlioptas teaches training, by the processing device, a large language model using the dataset to generate individual compact graphs from text inputs (Achlioptas, abstract, the paper describes the most extensive existing corpus of natural language utterances describing shape differences: ShapeTalk. ShapeTalk contains over half a million discriminative utterances produced by contrasting the shapes of common 3D objects for a variety of object classes and degrees of similarity. We also introduce a generic framework, ChangeIt3D, which builds on ShapeTalk and can use an arbitrary 3D generative model of shapes to produce edits that align the output better with the edit or deformation description. Finally, we introduce metrics for the quantitative evaluation of language-assisted shape editing methods that reflect key desiderata within this editing setup. We note that ShapeTalk allows methods to be trained with explicit 3D-to-language data, bypassing the necessity of “lifting” 2D to 3D using methods like neural rendering, as required by extant 2D image-language foundation models.
Page 3, col 2, par 2, … Specifically, we instruct the 2,161 annotators of ShapeTalk to provide descriptions that differentiate the two shapes within a context, and do so class-by-class. This latter design choice lessens their cognitive burden allowing them to transfer experience of annotating past recent examples within the same shape class. It is also worth noting that for each class, we provide visual examples of objects with annotations of part-names as well as names for different shape styles (e.g. “bowler” hat vs. “ivy” hat), without requiring their usage in the annotations. Typical resultant annotations of ShapeTalk can be seen in Figure 1.
Page 4, col 1, par 1, ShapeTalk utterances are highly diverse. To shed light
on the types of language used in ShapeTalk, we manually curate a large subset of word groups from user utterances into 5 different categories shown in Figure 2. We connect an utterance to these categories according to word membership. Note that in addition to these categories, an utterance can also contain “holistic” shape information if it does not reference any Parts or Local features, but rather describes the whole shape. Table 2 shows the proportion of utterances that contain different kinds of information, according to word membership.
Liu, abstract, the invention describes methods to segment, using at least one machine learning model, a first representation of an object into a plurality of parts according to a hierarchy, where the hierarchy indicates groups of datapoints for at least one part of the plurality of parts based on a value. The one or more processors generate at least one second representation of the object based on the at least one of the part of the plurality of parts.);
fine-tuning, by the processing device, the large language model using synthetic training data generated using a vision language model (Achlioptas, page 5, par 2-3, For modularity, we train the editing process via a two-stage approach. As the generative model G needs to capture sufficient geometric information, we pretrain an autoencoder during the first stage to achieve good reconstructions. Once G is pretrained, L is trained to associate higher utterance-compatibility with the target shape than with the distractor shape, using the latent representations given by the pretrained network G as input.
In the second stage of our approach, we link the frozen networks L and G together via the Shape Editor E, a low capacity network that learns to find editing latents in the space of G through predicting an update vector by regressing its magnitude and direction independently. This ‘update vector’ is then applied onto the source shape latent representation additively, promoting a direct interaction between the source representation and the underlying edit’s latent. During this stage, our model uses frozen weights for the encoder, decoder and neural listener L, and learns the weights for the shape editor E so as to 1) preserve similarity to the original input shape through regularization of the update magnitude and 2) maximize the L-evaluated utterance compatibility of the updated shape latent representation over that of the original shape.
Liu, [0032] approaches in accordance with various embodiments can be used to generate one or more parameters for a content generation environment. In at least one embodiment, a trained machine learning (ML) and/or artificial intelligence (AI) system, such as a large language model (LLM) or a vision language model (VLM), may be used to generate parameters for the content generation environment, such as, but not limited to, camera settings, scene lighting, video parameters, and/or the like, used for displaying objects within a scene. The parameters may be based on an input provided by a user or a proxy for a user to a trained language model (e.g., LLM, VLM, etc.) … For embodiments that incorporate one or more language models-that is, one or more LLMs, one or more VLMs, or a combination of LLMs and VLMs, the language model(s) may receive an input (e.g., a prompt, a request, a query, etc.) that is parsed or otherwise formatted to generate a deterministic output.);
Liu, Chen and Achlioptas are analogous art because they all teach method of AI generating image objects. Liu further teaches generating a hierarchical structure for multiple parts of the image object. Chen further teaches using Large Language Model to summarize text description from different dimensions and detect object segmentation of interest. Achlioptas further teaches using dataset to train LLM and VLM. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AI method of generating image object (taught in Liu and Chen), to further use the LLM and VLM (taught in Achlioptas), so as to provide convenient tool to help visually-impaired users to interact with object of interest and change them to better fit their design needs (Achlioptas, page 1, col 2, par 3).
The combination of Liu, Chen and Achlioptas further teaches integrating, by the processing device, the large language model with an interpreter that converts the individual compact graphs into corresponding object models (Liu, [0079] FIG. 9B is a flow diagram showing a method 910 for segmenting an object (e.g., 3D volume 300) into a plurality of parts (e.g., parts 136) and generating a representation of the object based on at least one part of the plurality of parts, in accordance with some implementations of the present disclosure. At block 912, the method 910 include segmenting a first representation of an object into a plurality of parts according to a hierarchy (e.g., hierarchical K-means clustering). The hierarchy can indicate groups of datapoints (e.g., points) for at least one part of the plurality of parts based on a value. The value can be a granularity value and can indicate a number of parts for which to segment the first representation of the object into.); and
generating or editing the corresponding object models based on object descriptions or edit descriptions received as inputs to the large language model (Achlioptas, abstract, the paper teaches a generic framework, ChangeIt3D, which builds on ShapeTalk and can use an arbitrary 3D generative model of shapes to produce edits that align the output better with the edit or deformation description. Finally, we introduce metrics for the quantitative evaluation of language-assisted shape editing methods that reflect key desiderata within this editing setup. We note that ShapeTalk allows methods to be trained with explicit 3D-to-language data, bypassing the necessity of “lifting” 2D to 3D using methods like neural rendering, as required by extant 2D image-language foundation models.
Page 8, col 1, par 2, … Figure 4 shows qualitative examples of decoded shape-edits with our ImNet-based models. These examples showcase the ability of ShapeTalk and ChangeIt3D to give rise to nuanced yet semantically correct edits to shapes. Moreover, our method appears to be able to preserve the overall identity of the input shape, and oftentimes create localized/minimal shape edits that can cover both structural and continuous changes.).
Regarding Claim 11. The combination of Liu, Chen and Achlioptas further teaches The method of claim 10, further comprising generating the synthetic training data by:
rendering multi-view images of three dimensional models (Liu, [0056] The system 200 may include the 3D representation generator 102, the feature space generator 110, the combiner 122, the dimensionality expander 126, and the feature field generator 130 as described with respect to the system 100. The 3D representation generator 102 may receive an input training geometry 202 indicative of a training 3D volume. The training 3D volume may be included in a training dataset. The training dataset may include 3D labels and 2D labels of a plurality of training 3D volumes. The 3D labels may be derived from an annotated shape dataset without textures with information on internal structures of the training 3D volumes. The 3D labels may be derived from an annotated shape dataset without textures with information on internal structures of the training 3D volumes. The 2D labels may be derived from multi-view rendering of the training 3D volumes by a 2D segmentation model. Prior to updating (e.g., training), the feature field generator 130 may include a plurality of priors and a plurality of weights which may be randomly selected. The plurality of priors can include 3D data, data from 2D foundation models, and heuristics such as convexity, geometric primitives, or color, among others. The plurality of weights may be iteratively updated during the updating process until the feature field generator 130 reaches convergence.);
captioning the multi-view images using the vision language model (Chen, abstract, the invention describes a method comprising the steps of obtaining a to-be-classified image, and determining a multidimensional text associated with the to-be-classified image; based on the texts of the multiple dimensions, determining text feature representations of the multiple dimensions; according to the multidimensional text feature representation, an extended text of the to-be-classified image is generated, and semantic information expressed by the extended text is more than semantic information expressed by the multi-dimensional text; extracting text semantic features from the extended text to obtain extended text semantic features; extracting image semantic features of the to-be-classified images; and based on the extended text semantic feature and the image semantic feature, classifying the to-be-classified image to obtain a category to which the to-be-classified image belongs. By adopting the method, the accuracy of image classification can be improved.
[0147], … Foreground features are extracted from the foreground region, subject features from the subject region, and background features from the background region. Based on the foreground, subject, and background features, descriptive text for the scene in the image to be classified is generated. This allows the generated descriptive text to describe the scene presented in the image from the perspective of the photographed subject, foreground, and background, and to describe the image scene from the overall layout, thus obtaining a macroscopic description of the scene.
Liu, [0032] approaches in accordance with various embodiments can be used to generate one or more parameters for a content generation environment. In at least one embodiment, a trained machine learning (ML) and/or artificial intelligence (AI) system, such as a large language model (LLM) or a vision language model (VLM), may be used to generate parameters for the content generation environment, such as, but not limited to, camera settings, scene lighting, video parameters, and/or the like, used for displaying objects within a scene. The parameters may be based on an input provided by a user or a proxy for a user to a trained language model (e.g., LLM, VLM, etc.) … For embodiments that incorporate one or more language models-that is, one or more LLMs, one or more VLMs, or a combination of LLMs and VLMs, the language model(s) may receive an input (e.g., a prompt, a request, a query, etc.) that is parsed or otherwise formatted to generate a deterministic output.); and
generating text descriptions corresponding to the captions using a separate large language model (Chen, [n0224] Image semantic structuring generates text describing objects or scenes in an image from different perspectives, i.e., multi-dimensional text. However, the semantic information expressed by this text is relatively fragmented and is not presented in the form of natural language. Therefore, in the LLM structured semantic induction part, this embodiment further uses a large language model to summarize the above-mentioned scattered semantic information, thereby generating a text description in natural language form.).
The reasoning for combination of Liu, Chen and Achlioptas is the same as described in Claim 10.
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Liu et al (US20260245300) in view of Chen et al (CN117253237), Achlioptas et al. further in view of Yin et al (US20240290054).
Regarding Claim 12. The combination of Liu, Chen and Achlioptas fails to explicitly teach, however, Yin teaches The method of claim 11, wherein the vision language model comprises a pre-trained Contrastive Language Image Pre-training neural network model (Yin, abstract, the invention describes generation method for three-dimensional (3D) object models for users without a sufficient skill set for content creation. One or more style transfer networks may be combined with a generative network to generate objects based on parameters associated with a textual input. An input including a 3D mesh and texture may be provided to a trained system along with a textual input that includes parameters for object generation. Features of the input object may be identified and then tuned in accordance with the textual input to generate a modified 3D object that includes a new texture along with one or more geometric adjustments.
[0015]… In at least one embodiment, a pretrained language-vision model, such as contrastive language-image pre-training (CLIP), produces a joint embedding of image and text. Using CLIP, systems and methods may render result into images, obtain embeddings of the rendered images, and try to match the embedding of the input text. The system may be improved by evaluating different costs or losses, where the costs or losses may be tuned to provide more preference to the input 3D shape or to the text input. Costs or losses may include style costs or losses, content costs or losses, or CLIP costs or losses, as examples. Various embodiments enable production of novel 3D assets as stylization results.), and the captioning the multi-view images comprises:
generating image embeddings for the multi-view images using the pre-trained Contrastive Language Image Pre-training neural network model (Yin, [0015]… In at least one embodiment, a pretrained language-vision model, such as contrastive language-image pre-training (CLIP), produces a joint embedding of image and text. Using CLIP, systems and methods may render result into images, obtain embeddings of the rendered images, and try to match the embedding of the input text. The system may be improved by evaluating different costs or losses, where the costs or losses may be tuned to provide more preference to the input 3D shape or to the text input. Costs or losses may include style costs or losses, content costs or losses, or CLIP costs or losses, as examples. Various embodiments enable production of novel 3D assets as stylization results.
[0017] … In various embodiments, camera data 114 corresponds to camera extrinsic information to enable generation of different 3D outputs from a variety of different view directions. Using the input camera data 114, the ccGAN 104 may include a viewpoint conditioned GAN which can be used to produce text-guided stylized multi-view images in combination with the 3D stylization module 106 to edit a 3D mesh and texture based on the stylized multi-view imagery.);
generating text embeddings for a set of candidate captions using the pre-trained Contrastive Language Image Pre-training neural network model (Yin, [0025] FIG. 2 illustrates an environment 200 representative of an architecture of a generative neural network, such as ccGAN 104. … In this example, the ccGAN 104 includes a generator (G) 202 and a discriminator (D) 204 where the generator 202 generates one or more outputs 206 for review and analysis by the discriminator 204 to make a determination 208 of whether or not the output 206 is in a particular view. That is, the discriminator 204 determines whether an image produced by the generator 202 corresponds to a desired view.
[0027] Further illustrated is a camera view 222 that is used to generate a conditional vector 224 of the camera view (u), associated with one or more camera extrinsics. The vector 224 is passed to both the generator 202 and discriminator 204 such that the generator 202 is able to learn to generate images in the specific view and the discriminator 204 is able to tell whether the generated images are in the right (e.g., desired) views. As discussed herein, u is embedded using a multilayer perceptron (MLP) 226 and the resulting output embedding 228 (embed(u)) is concatenated with the vector (w) 216, which may further be processed by another MLP 230.
[0031] Various embodiments may further incorporate a stylized generator (G') to produce images in unseen styles using guidance from a language-vision model. These techniques may be applied to the generator 202 to produce a generator that is able to generate multi-view consistent renders of 3D objects in a specific style, unseen in the training data. For example, the training may be performed for a fixed style, such as "zombie animal" or "cartoon features" using guidance from one or more pretrained language- vision (e.g., CLIP) models. As noted herein, random sample noise 218 and the camera views 222 may be used to drive a frozen generator and trainable generator to synthesize images in the same random identity and random view. Accordingly, directional CLIP loss may be used to learn domain shifts when incorporating G'.); and
selecting captions for the multi-view images based on similarities between the image embeddings and text embeddings (Yin, [0015]… In at least one embodiment, a pretrained language-vision model, such as contrastive language-image pre-training (CLIP), produces a joint embedding of image and text. Using CLIP, systems and methods may render result into images, obtain embeddings of the rendered images, and try to match the embedding of the input text. The system may be improved by evaluating different costs or losses, where the costs or losses may be tuned to provide more preference to the input 3D shape or to the text input. Costs or losses may include style costs or losses, content costs or losses, or CLIP costs or losses, as examples. Various embodiments enable production of novel 3D assets as stylization results.
[0032] FIG. 3 illustrates an environment 300 representing the 3D stylization module 106 along with a generative neural network (e.g., ccGAN 104) to generate one or more losses that may be used to generate one or more images based on the image input 110 and the text input 112. In this example, the ccGAN 104 may include one or more features of FIG. 2 and, moreover, may incorporate the stylization generator in order to generate one or more outputs for evaluation by the discriminator in order to determine a loss or cost. In various embodiments, the stylization generator produces multi-view images in an unseen text-driven style that may be used to help train the 3D stylization module 106 to modify a geometry and texture of input 3D models to a style. After training, the module can be used to style new 3D meshes using the style in real or near-real time (e.g., without significant delay).).
Liu, Chen, Achlioptas and Yin are analogous art because they all teach method of AI generating image objects. The combination of Liu, Chen and Achlioptas further teaches generating a hierarchical structure for multiple parts of the image object and using dataset to train LLM and VLM to manipulate the hierarchical structure. Yin further teaches training a multi-view style transfer neural network that can generate objects based on parameters associated with a text input. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AI method of generating image object with hierarchical parts structure (taught in Liu, Chen and Achlioptas), to further use the multi-view style transfer neural network that can generate objects based on parameters associated with a text input (taught in Yin), so as to provide use with a model that can styling 3D object and generate realistic look without significant levels of skill and large amount of time (Yin, [0001]).
Claims 13 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al (US20260245300) in view of Chen et al (CN117253237), Achlioptas et al. further in view of Koo et al. ("Partglot: Learning shape part segmentation from language reference games.", 2022).
Regarding Claim 13. The combination of Liu, Chen and Achlioptas further teaches The method of claim 10, wherein the training includes:
obtaining transforms from three dimensional mesh components of the object models that capture spatial relationships and hierarchical structures (Achlioptas, Page 8, col 1, par 2, … Figure 4 shows qualitative examples of decoded shape-edits with our ImNet-based models. These examples showcase the ability of ShapeTalk and ChangeIt3D to give rise to nuanced yet semantically correct edits to shapes. Moreover, our method appears to be able to preserve the overall identity of the input shape, and oftentimes create localized/minimal shape edits that can cover both structural and continuous changes.
Liu, [0029] The systems and methods in accordance with the present disclosure may be applied to various technical fields, such as 3D shape generation, simulation, autonomous driving, animation, and editing AI-generated meshes. For example, systems and methods in accordance with the present disclosure enable generation of meshes per part instead of an entire object. In some implementations, the systems and methods described herein may provide segmentation data to a simulation environment (e.g., NVIDIA's DriveSIM) for the simulation environment to identify regions of interest (e.g., legs, arms) to perform operations associated with the segments of the 3D part. The simulation environment may be built into applications or software platforms for creating, generating, modifying, or manipulating 3D models or parts.);
generating procedural compact graphs based on the transforms (Liu, [0029], … The simulation environment may be built into applications or software platforms for creating, generating, modifying, or manipulating 3D models or parts.
[0052] As an example, responsive to the clustering algorithm being agglomerative clustering, each data point (e.g., point) on the part feature field 132 is treated as an individual cluster. The distance between each pair of clusters is then calculated. A computation of all the distances of each pair of clusters in the part feature field 132 may be represented by a distance matrix. The distance may be calculated by, for example, Euclidean or Manhattan distance. Based on the distance, the two clusters that are closest to each other are merged together to reduce a number of clusters of the part feature field 132. After merging the two clusters, the distance matrix can be updated to reflect distances between the merged cluster and remaining clusters. To compute the distances between these clusters, the feature field part generator 134 may use at least one of single linkage, complete linkage, average linkage, or Ward's method. Merging two clusters and updating the distance matrix is then repeated until either all data points on the part feature field 132 are in a single cluster, or until a threshold is met. The threshold can be a predefined number of clusters (e.g., predefined by the user). A hierarchical tree (e.g., dendrogram) can also be formed based on the clustering process. The hierarchical tree can be cut at a specified level to obtain a desired number of clusters.
[0081] At block 914, the method 910 includes generating at least one second representation of the object based on at least one part of the plurality of parts. The first representation of the object and the at least one second representation of the object can be 3D. In various implementations, the method 910 can include adjusting the at least one second representation according to the at least one part of the plurality of parts. For example, the at least one second representation can be edited following generation using the at least one of the plurality of parts.
Therefore, the hierarchical tree of parts of the object can be changed by updating the distance matrix and/or clustering.); and
The combination of Liu, Chen and Achlioptas fails to explicitly teach, however, Koo teaches training the large language model to map between the text descriptions and the procedural compact graphs (Koo, abstract, the paper teaches PartGlot, a neural framework and associated architectures for learning semantic part segmentation of 3D shape geometry, based solely on part referential language. We exploit the fact that linguistic descriptions of a shape can provide priors on the shape’s parts – as natural language has evolved to reflect human perception of the compositional structure of objects, essential to their recognition and use. For training we use ShapeGlot’s paired geometry / language data collected via a reference game where a speaker produces an utterance to differentiate a target shape from two distractors and the listener has to find the target based on this utterance [3]. Our network is designed to solve this target multi-modal recognition problem, by carefully incorporating a Transformer-based attention module so that the output attention can precisely highlight the semantic part or parts described in the language. Remarkably, the network operates without any direct supervision on the 3D geometry itself. Furthermore, we also demonstrate that the learned part information is generalizable to shape classes unseen during training. Our approach opens the possibility of learning 3D shape parts from language alone, without the need for large-scale part geometry annotations, thus facilitating annotation acquisition.).
Liu, Chen, Achlioptas and Koo are analogous art because they all teach method of AI generating image objects. The combination of Liu, Chen and Achlioptas further teaches generating a hierarchical structure for multiple parts of the image object and using dataset to train LLM and VLM to manipulate the hierarchical structure. Koo further teaches training a neural framework for learning semantic part segmentation of 3D shape geometry based on part referential language. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AI method of generating image object with hierarchical parts structure (taught in Liu, Chen and Achlioptas), to further training a neural framework for learning semantic part segmentation of 3D shape geometry based on part referential language (taught in Koo), so as to provide neural framework allowing learning 3D shape parts without need for large-scale part geometry annotations (Koo, page 16505, col 2, par 1).
Claims 14 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al (US20260245300) in view of Chen et al (CN117253237), Achlioptas et al. further in view of Tian et al. ("Shapescaffolder: Structure-aware 3d shape generation from text.", IEEE, 2023.).
Regarding Claim 14. The combination of Liu, Chen and Achlioptas fails to explicitly teach, however, Tian teaches The method of claim 10, further comprising:
receiving a text prompt describing a three dimensional object (Tian, abstract, the paper describes ShapeScaffolder, a structure-based neural network for generating colored 3D shapes based on text input. The approach, similar to providing scaffolds as internal structural supports and adding more details to them, aims to capture finer text-shape connections and improve the quality of generated shapes. … We first build the structured shape implicit fields in an unsupervised manner. We then propose the part-level attention mechanism between shape parts and textual graph nodes to align the two modalities at the structural level. Finally, we employ a shape refiner to add further detail to the predicted structure, yielding the final results. Extensive experimentation demonstrates that our approaches outperform state-of-the-art methods in terms of both shape fidelity and shape-text matching. Our methods also allow for part level manipulation and improved part-level completeness.
Figure 1, input text is “brown wooden chair with curved armrests and a blue plush seat cushion”, the structure graph displays a set of parts of the 3D chair object.);
Liu, Chen, Achlioptas and Tian are analogous art because they all teach method of AI generating image objects. The combination of Liu, Chen and Achlioptas further teaches generating a hierarchical structure for multiple parts of the image object and using dataset to train LLM and VLM to manipulate the hierarchical structure. Tian further teaches a structure-based neural network for generating colored 3D shapes based on text input with shape refiner for adding further detail to the structure. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AI method of generating image object with hierarchical parts structure (taught in Liu, Chen and Achlioptas), to further use the structure-based neural network for generating colored 3D shapes based on text input with shape refiner for adding further detail to the structure (taught in Tian), so as to provide method to generate 3D shape with shape fidelity and shape-text matching (Tian, abstract).
The combination of Liu, Chen, Achlioptas and Tian further teaches generating, using the large language model, a compact graph representing the described three-dimensional object (Liu, abstract, the invention describes methods to segment, using at least one machine learning model, a first representation of an object into a plurality of parts according to a hierarchy, where the hierarchy indicates groups of datapoints for at least one part of the plurality of parts based on a value. The one or more processors generate at least one second representation of the object based on the at least one of the part of the plurality of parts.
[0032] Approaches in accordance with various embodiments can be used to generate one or more parameters for a content generation environment. In at least one embodiment, a trained machine learning (ML) and/or artificial intelligence (AI) system, such as a large language model (LLM) or a vision language model (VLM), may be used to generate parameters for the content generation environment, such as, but not limited to, camera settings, scene lighting, video parameters, and/or the like, used for displaying objects within a scene. The parameters may be based on an input provided by a user or a proxy for a user to a trained language model (e.g., LLM, VLM, etc.) that can then generate one or more settings in accordance with the input.);
converting, using the interpreter, the compact graph into a three-dimensional mesh object (Liu, [0029] The systems and methods in accordance with the present disclosure may be applied to various technical fields, such as 3D shape generation, simulation, autonomous driving, animation, and editing AI-generated meshes. For example, systems and methods in accordance with the present disclosure enable generation of meshes per part instead of an entire object. In some implementations, the systems and methods described herein may provide segmentation data to a simulation environment (e.g., NVIDIA's DriveSIM) for the simulation environment to identify regions of interest (e.g., legs, arms) to perform operations associated with the segments of the 3D part. The simulation environment may be built into applications or software platforms for creating, generating, modifying, or manipulating 3D models or parts.); and
rendering an image of the three dimensional mesh for display (Liu, [0037] Once the 3D representation generator 102 receives the input geometry 104, the 3D representation generator 102 can generate at least one 3D representation. The 3D representation can model and represent the geometry of the 3D volume. The 3D representation can have different formats such as, but not limited to, a signed distance function (SDF) volume (e.g., encoded distance for each point of the input geometry 104), a point cloud (e.g., collection of spatial coordinates of the input geometry 104), a mesh representation (e.g., including vertices, edges, faces), a volumetric pixel (e.g., voxel) grid, or a parameter point model.).
Regarding Claim 15. The combination of Liu, Chen, Achlioptas and Tian further teaches The method of claim 14, further comprising:
receiving a text edit instruction describing a modification to the three-dimensional object (Achlioptas, Page 8, col 1, par 2, … Figure 4 shows qualitative examples of decoded shape-edits with our ImNet-based models. These examples showcase the ability of ShapeTalk and ChangeIt3D to give rise to nuanced yet semantically correct edits to shapes. Moreover, our method appears to be able to preserve the overall identity of the input shape, and oftentimes create localized/minimal shape edits that can cover both structural and continuous changes.);
modifying, using the large language model, the compact graph based on the text edit instruction (Tian, page 8, col 2, par 1, Part-level Manipulation: The ability to learn structure-based 3D shape representations and align them with textual descriptions enables precise and convenient manipulation of shapes. In the process of generating a shape from text, our method allows for manipulation at various levels, including modification of the part-level structure, attributes, or colors. The results of this manipulation are illustrated in Fig. 7. The changes made in the text appear to reflect the changes in part structure (a), part attribute (c), and part-related color refinement (b), suggesting a possible consistency between the constituent parts of the shape and the corresponding text.
Liu, [0032] Approaches in accordance with various embodiments can be used to generate one or more parameters for a content generation environment. In at least one embodiment, a trained machine learning (ML) and/or artificial intelligence (AI) system, such as a large language model (LLM) or a vision language model (VLM), may be used to generate parameters for the content generation environment, such as, but not limited to, camera settings, scene lighting, video parameters, and/or the like, used for displaying objects within a scene. The parameters may be based on an input provided by a user or a proxy for a user to a trained language model (e.g., LLM, VLM, etc.) that can then generate one or more settings in accordance with the input.);
converting, using the interpreter, the modified compact graph into an updated three dimensional mesh (Liu, [0029] The systems and methods in accordance with the present disclosure may be applied to various technical fields, such as 3D shape generation, simulation, autonomous driving, animation, and editing AI-generated meshes. For example, systems and methods in accordance with the present disclosure enable generation of meshes per part instead of an entire object. In some implementations, the systems and methods described herein may provide segmentation data to a simulation environment (e.g., NVIDIA's DriveSIM) for the simulation environment to identify regions of interest (e.g., legs, arms) to perform operations associated with the segments of the 3D part. The simulation environment may be built into applications or software platforms for creating, generating, modifying, or manipulating 3D models or parts.); and
rendering an updated image of the updated three dimensional mesh for display (Liu, [0037] Once the 3D representation generator 102 receives the input geometry 104, the 3D representation generator 102 can generate at least one 3D representation. The 3D representation can model and represent the geometry of the 3D volume. The 3D representation can have different formats such as, but not limited to, a signed distance function (SDF) volume (e.g., encoded distance for each point of the input geometry 104), a point cloud (e.g., collection of spatial coordinates of the input geometry 104), a mesh representation (e.g., including vertices, edges, faces), a volumetric pixel (e.g., voxel) grid, or a parameter point model.).
The reasoning for combination of Liu, Chen, Achlioptas and Tian is the same as described in Claim 14.
Claims 16, 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al (US20260245300) in view of Tian et al. ("Shapescaffolder: Structure-aware 3d shape generation from text.", IEEE, 2023.).
Regarding Claim 16. Liu teaches A system comprising:
a memory component; and
one or more processing devices coupled to the memory component, the processing devices operable (Liu, abstract, the invention describes methods to segment, using at least one machine learning model, a first representation of an object into a plurality of parts according to a hierarchy, where the hierarchy indicates groups of datapoints for at least one part of the plurality of parts based on a value. The one or more processors generate at least one second representation of the object based on the at least one of the part of the plurality of parts.
[0055] With reference to FIG. 2, FIG. 2 is an example system 200 for updating a system to segment a 3D volume, in accordance with some implementations of the present disclosure. … For instance, various functions may be carried out by a processor executing instructions stored in memory.) to:
create an initial compact graph representing an initial hierarchy of initial object attributes by inputting an object description into a large language model (Liu, abstract, the invention describes methods to segment, using at least one machine learning model, a first representation of an object into a plurality of parts according to a hierarchy, where the hierarchy indicates groups of datapoints for at least one part of the plurality of parts based on a value. The one or more processors generate at least one second representation of the object based on the at least one of the part of the plurality of parts.
[0079] FIG. 9B is a flow diagram showing a method 910 for segmenting an object (e.g., 3D volume 300) into a plurality of parts (e.g., parts 136) and generating a representation of the object based on at least one part of the plurality of parts, in accordance with some implementations of the present disclosure. At block 912, the method 910 include segmenting a first representation of an object into a plurality of parts according to a hierarchy (e.g., hierarchical K-means clustering). The hierarchy can indicate groups of datapoints (e.g., points) for at least one part of the plurality of parts based on a value. The value can be a granularity value and can indicate a number of parts for which to segment the first representation of the object into.)
Liu fails to explicitly teach, however Tian teaches an object description (Tian, abstract, the paper describes ShapeScaffolder, a structure-based neural network for generating colored 3D shapes based on text input. The approach, similar to providing scaffolds as internal structural supports and adding more details to them, aims to capture finer text-shape connections and improve the quality of generated shapes. … We first build the structured shape implicit fields in an unsupervised manner. We then propose the part-level attention mechanism between shape parts and textual graph nodes to align the two modalities at the structural level. Finally, we employ a shape refiner to add further detail to the predicted structure, yielding the final results. Extensive experimentation demonstrates that our approaches outperform state-of-the-art methods in terms of both shape fidelity and shape-text matching. Our methods also allow for part level manipulation and improved part-level completeness.
Figure 1, input text is “brown wooden chair with curved armrests and a blue plush seat cushion”, the structure graph displays a set of parts of the 3D chair object. The text graph displays a hierarchical structure of object attributes.);
Liu and Tian are analogous art because they both teach method of AI generating image objects. Liu further teaches generating a hierarchical structure for multiple parts of the image object and using dataset to train LLM and VLM to manipulate the hierarchical structure. Tian further teaches a structure-based neural network for generating colored 3D shapes based on text input with shape refiner for adding further detail to the structure. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AI method of generating image object with hierarchical parts structure (taught in Liu), to further use the structure-based neural network for generating colored 3D shapes based on text input with shape refiner for adding further detail to the structure (taught in Tian), so as to provide method to generate 3D shape with shape fidelity and shape-text matching (Tian, abstract).
The combination of Liu and Tian further teaches generate an initial object model based on the initial compact graph by applying one or more initial geometric transformations that convert the initial compact graph into an initial three dimensional mesh representation (Liu, [0029] The systems and methods in accordance with the present disclosure may be applied to various technical fields, such as 3D shape generation, simulation, autonomous driving, animation, and editing AI-generated meshes. For example, systems and methods in accordance with the present disclosure enable generation of meshes per part instead of an entire object. In some implementations, the systems and methods described herein may provide segmentation data to a simulation environment (e.g., NVIDIA's DriveSIM) for the simulation environment to identify regions of interest (e.g., legs, arms) to perform operations associated with the segments of the 3D part. The simulation environment may be built into applications or software platforms for creating, generating, modifying, or manipulating 3D models or parts.
[0037] Once the 3D representation generator 102 receives the input geometry 104, the 3D representation generator 102 can generate at least one 3D representation. The 3D representation can model and represent the geometry of the 3D volume. The 3D representation can have different formats such as, but not limited to, a signed distance function (SDF) volume (e.g., encoded distance for each point of the input geometry 104), a point cloud (e.g., collection of spatial coordinates of the input geometry 104), a mesh representation (e.g., including vertices, edges, faces), a volumetric pixel (e.g., voxel) grid, or a parameter point model.);
create an updated compact graph representing an updated hierarchy of updated object attributes by inputting an object edit into the large language model (Tian, page 8, col 2, par 1, Part-level Manipulation: The ability to learn structure-based 3D shape representations and align them with textual descriptions enables precise and convenient manipulation of shapes. In the process of generating a shape from text, our method allows for manipulation at various levels, including modification of the part-level structure, attributes, or colors. The results of this manipulation are illustrated in Fig. 7. The changes made in the text appear to reflect the changes in part structure (a), part attribute (c), and part-related color refinement (b), suggesting a possible consistency between the constituent parts of the shape and the corresponding text.); and
generate an updated object model based on the updated compact graph by applying one or more updated geometric transformations that convert the updated compact graph into an updated three dimensional mesh representation (Liu, [0029] The systems and methods in accordance with the present disclosure may be applied to various technical fields, such as 3D shape generation, simulation, autonomous driving, animation, and editing AI-generated meshes. For example, systems and methods in accordance with the present disclosure enable generation of meshes per part instead of an entire object. In some implementations, the systems and methods described herein may provide segmentation data to a simulation environment (e.g., NVIDIA's DriveSIM) for the simulation environment to identify regions of interest (e.g., legs, arms) to perform operations associated with the segments of the 3D part. The simulation environment may be built into applications or software platforms for creating, generating, modifying, or manipulating 3D models or parts.
[0037] Once the 3D representation generator 102 receives the input geometry 104, the 3D representation generator 102 can generate at least one 3D representation. The 3D representation can model and represent the geometry of the 3D volume. The 3D representation can have different formats such as, but not limited to, a signed distance function (SDF) volume (e.g., encoded distance for each point of the input geometry 104), a point cloud (e.g., collection of spatial coordinates of the input geometry 104), a mesh representation (e.g., including vertices, edges, faces), a volumetric pixel (e.g., voxel) grid, or a parameter point model.).
Regarding Claim 18. The combination of Liu and Tian further teaches The system of claim 16, wherein the initial compact graph and the updated compact graph comprise nodes representing geometric primitives and operations for combining and modifying the geometric primitives (Liu, [0052] As an example, responsive to the clustering algorithm being agglomerative clustering, each data point (e.g., point) on the part feature field 132 is treated as an individual cluster. The distance between each pair of clusters is then calculated. A computation of all the distances of each pair of clusters in the part feature field 132 may be represented by a distance matrix. The distance may be calculated by, for example, Euclidean or Manhattan distance. Based on the distance, the two clusters that are closest to each other are merged together to reduce a number of clusters of the part feature field 132. After merging the two clusters, the distance matrix can be updated to reflect distances between the merged cluster and remaining clusters. To compute the distances between these clusters, the feature field part generator 134 may use at least one of single linkage, complete linkage, average linkage, or Ward's method. Merging two clusters and updating the distance matrix is then repeated until either all data points on the part feature field 132 are in a single cluster, or until a threshold is met. The threshold can be a predefined number of clusters (e.g., predefined by the user). A hierarchical tree (e.g., dendrogram) can also be formed based on the clustering process. The hierarchical tree can be cut at a specified level to obtain a desired number of clusters.
[0056] The system 200 may include the 3D representation generator 102, the feature space generator 110, the combiner 122, the dimensionality expander 126, and the feature field generator 130 as described with respect to the system 100. The 3D representation generator 102 may receive an input training geometry 202 indicative of a training 3D volume. The training 3D volume may be included in a training dataset. The training dataset may include 3D labels and 2D labels of a plurality of training 3D volumes. The 3D labels may be derived from an annotated shape dataset without textures with information on internal structures of the training 3D volumes. The 3D labels may be derived from an annotated shape dataset without textures with information on internal structures of the training 3D volumes. The 2D labels may be derived from multi-view rendering of the training 3D volumes by a 2D segmentation model. Prior to updating (e.g., training), the feature field generator 130 may include a plurality of priors and a plurality of weights which may be randomly selected. The plurality of priors can include 3D data, data from 2D foundation models, and heuristics such as convexity, geometric primitives, or color, among others. The plurality of weights may be iteratively updated during the updating process until the feature field generator 130 reaches convergence.).
Regarding Claim 19. The combination of Liu and Tian further teaches The system of claim 18, wherein the nodes include at least one of: a cylinder node, a rectangle node, a point instance node, a transform node, a fillet node, a fill node, an extrude node, and a join node (Liu, [0056] The system 200 may include the 3D representation generator 102, the feature space generator 110, the combiner 122, the dimensionality expander 126, and the feature field generator 130 as described with respect to the system 100. The 3D representation generator 102 may receive an input training geometry 202 indicative of a training 3D volume. The training 3D volume may be included in a training dataset. The training dataset may include 3D labels and 2D labels of a plurality of training 3D volumes. The 3D labels may be derived from an annotated shape dataset without textures with information on internal structures of the training 3D volumes. The 2D labels may be derived from multi-view rendering of the training 3D volumes by a 2D segmentation model. Prior to updating (e.g., training), the feature field generator 130 may include a plurality of priors and a plurality of weights which may be randomly selected. The plurality of priors can include 3D data, data from 2D foundation models, and heuristics such as convexity, geometric primitives, or color, among others. The plurality of weights may be iteratively updated during the updating process until the feature field generator 130 reaches convergence.
Tian, Figure 5 shows parts with primitive shape such as rectangle.).
Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Liu et al (US20260245300) in view of Tian et al. ("Shapescaffolder: Structure-aware 3d shape generation from text.", IEEE, 2023.), Chen et al (CN117253237) further in view of Achlioptas et al. ("ShapeTalk: A language dataset and framework for 3d shape edits and deformations." 2023 IEEE).
Regarding Claim 17. The combination of Liu and Tian further teaches The system of claim 16, wherein the processing devices are further operable to:
generate synthetic training data by rendering multi-view images of three dimensional models, (Liu, [0056] The system 200 may include the 3D representation generator 102, the feature space generator 110, the combiner 122, the dimensionality expander 126, and the feature field generator 130 as described with respect to the system 100. The 3D representation generator 102 may receive an input training geometry 202 indicative of a training 3D volume. The training 3D volume may be included in a training dataset. The training dataset may include 3D labels and 2D labels of a plurality of training 3D volumes. The 3D labels may be derived from an annotated shape dataset without textures with information on internal structures of the training 3D volumes. The 3D labels may be derived from an annotated shape dataset without textures with information on internal structures of the training 3D volumes. The 2D labels may be derived from multi-view rendering of the training 3D volumes by a 2D segmentation model. Prior to updating (e.g., training), the feature field generator 130 may include a plurality of priors and a plurality of weights which may be randomly selected. The plurality of priors can include 3D data, data from 2D foundation models, and heuristics such as convexity, geometric primitives, or color, among others. The plurality of weights may be iteratively updated during the updating process until the feature field generator 130 reaches convergence.); and
The combination of Liu and Tian fails to explicitly teach, however, Chen teaches captioning the multi-view images using a vision language model, and generating text descriptions corresponding to the captions using a separate language model (Chen, abstract, the invention describes a method comprising the steps of obtaining a to-be-classified image, and determining a multidimensional text associated with the to-be-classified image; based on the texts of the multiple dimensions, determining text feature representations of the multiple dimensions; according to the multidimensional text feature representation, an extended text of the to-be-classified image is generated, and semantic information expressed by the extended text is more than semantic information expressed by the multi-dimensional text; extracting text semantic features from the extended text to obtain extended text semantic features; extracting image semantic features of the to-be-classified images; and based on the extended text semantic feature and the image semantic feature, classifying the to-be-classified image to obtain a category to which the to-be-classified image belongs. By adopting the method, the accuracy of image classification can be improved.
[0147], … Foreground features are extracted from the foreground region, subject features from the subject region, and background features from the background region. Based on the foreground, subject, and background features, descriptive text for the scene in the image to be classified is generated. This allows the generated descriptive text to describe the scene presented in the image from the perspective of the photographed subject, foreground, and background, and to describe the image scene from the overall layout, thus obtaining a macroscopic description of the scene.
[n0224] Image semantic structuring generates text describing objects or scenes in an image from different perspectives, i.e., multi-dimensional text. However, the semantic information expressed by this text is relatively fragmented and is not presented in the form of natural language. Therefore, in the LLM structured semantic induction part, this embodiment further uses a large language model to summarize the above-mentioned scattered semantic information, thereby generating a text description in natural language form.
Liu, [0032] approaches in accordance with various embodiments can be used to generate one or more parameters for a content generation environment. In at least one embodiment, a trained machine learning (ML) and/or artificial intelligence (AI) system, such as a large language model (LLM) or a vision language model (VLM), may be used to generate parameters for the content generation environment, such as, but not limited to, camera settings, scene lighting, video parameters, and/or the like, used for displaying objects within a scene. The parameters may be based on an input provided by a user or a proxy for a user to a trained language model (e.g., LLM, VLM, etc.) … For embodiments that incorporate one or more language models-that is, one or more LLMs, one or more VLMs, or a combination of LLMs and VLMs, the language model(s) may receive an input (e.g., a prompt, a request, a query, etc.) that is parsed or otherwise formatted to generate a deterministic output.).
The combination of Liu, Tian and Chen fails to explicitly teach, however, Achlioptas teaches fine-tune the large language model using the synthetic training data to improve performance on out of distribution object categories (Achlioptas, abstract, the paper describes the most extensive existing corpus of natural language utterances describing shape differences: ShapeTalk. ShapeTalk contains over half a million discriminative utterances produced by contrasting the shapes of common 3D objects for a variety of object classes and degrees of similarity. We also introduce a generic framework, ChangeIt3D, which builds on ShapeTalk and can use an arbitrary 3D generative model of shapes to produce edits that align the output better with the edit or deformation description. Finally, we introduce metrics for the quantitative evaluation of language-assisted shape editing methods that reflect key desiderata within this editing setup. We note that ShapeTalk allows methods to be trained with explicit 3D-to-language data, bypassing the necessity of “lifting” 2D to 3D using methods like neural rendering, as required by extant 2D image-language foundation models.
Page 5, par 2-3, For modularity, we train the editing process via a two-stage approach. As the generative model G needs to capture sufficient geometric information, we pretrain an autoencoder during the first stage to achieve good reconstructions. Once G is pretrained, L is trained to associate higher utterance-compatibility with the target shape than with the distractor shape, using the latent representations given by the pretrained network G as input.
In the second stage of our approach, we link the frozen networks L and G together via the Shape Editor E, a low capacity network that learns to find editing latents in the space of G through predicting an update vector by regressing its magnitude and direction independently. This ‘update vector’ is then applied onto the source shape latent representation additively, promoting a direct interaction between the source representation and the underlying edit’s latent. During this stage, our model uses frozen weights for the encoder, decoder and neural listener L, and learns the weights for the shape editor E so as to 1) preserve similarity to the original input shape through regularization of the update magnitude and 2) maximize the L-evaluated utterance compatibility of the updated shape latent representation over that of the original shape.);
Liu, Tian, Chen and Achlioptas are analogous art because they all teach method of AI generating image objects. Liu further teaches generating a hierarchical structure for multiple parts of the image object. Chen further teaches using Large Language Model to summarize text description from different dimensions and detect object segmentation of interest. Achlioptas further teaches using dataset to train LLM and VLM. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AI method of generating image object (taught in Liu, Tian and Chen), to further use the LLM and VLM (taught in Achlioptas), so as to provide convenient tool to help visually-impaired users to interact with object of interest and change them to better fit their design needs (Achlioptas, page 1, col 2, par 3).
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Liu et al (US20260245300) in view of Tian et al. ("Shapescaffolder: Structure-aware 3d shape generation from text.", IEEE, 2023.) further in view of Hizmi (“LLMto3D Generating parametric Objects using Large Language Model”, 2024, ACADIA).
Regarding Claim 20. The combination of Liu and Tian fails to explicitly teach, however, Hizmi teaches The system of claim 16, wherein the processing devices are further operable to:
display a user interface comprising an object preview window and a set of parameter controls;
update the set of parameter controls based on adjustable attributes defined in the initial compact graph; and
interpret user inputs received via the set of parameter controls as object edits for modifying the initial object model (Hizmi, abstract, the paper describes methods of using Machine Learning (ML) to generate 3D objects from textual descriptions. We introduce a novel method for translating natural language descrip-tions into parametric 3D objects using Large Language Models (LLMs). Our approach employs multiple agents, each one an LLM pre-trained for a specific task. The first agent deconstructs textual prompts into design elements and describes their geometry and spatial relations. The second agent translates the description into code using the Rhino. Geometry coding library in the Rhino3D-Grasshopper modeling environment. A final agent reassembles the models and adds parametric control interfaces, enabling customizable outputs. In this paper, we describe the method's architecture, and the training methodol-ogies used to fine-tune the models.
Page 7, col 2, par 2, Code Output: The output of the codes in the example dataset is a Python code describing a BREP geometry, which is developed by the LLM agents. This code is then transmitted back to the Grasshopper environment which realizes the geometry via the Hops components. In addition to the geometry, the process automatically generates input sliders through an interface invoked by the Python scripts and executed by the Hops components. These sliders enable users to tailor the final object precisely. An illustrative example demonstrates a kettle produced by this method (Figure 8). The output parameters allow for extensive manipulation to achieve the desired design and functionality of the kettle.).
Liu, Tian and Hizmi are analogous art because they all teach method of AI generating image objects. The combination of Liu and Tian further teaches generating a hierarchical structure for multiple parts of the image object using LLM/VLM. Hizmi further teaches parameter control interface for user to edit the 3D object model. Therefore, it would have been obvious to a person with ordinary skill in the art before the effective filing date of the claimed invention, to modify the AI method of generating image object (taught in Liu and Tian), to further use the parameter control interface (taught in Hizmi), so as to provide intuitive user interface for use to further customize 3D object for printability or manufacturability (Hizmi, abstract).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Cetintas et al (US20250299485), abstract, the invention teaches dynamic novel view reconstruction based at least in part on flow rematching. A first computing system can update a graph neural network based at least on video data representing a plurality of first objects and a plurality of first labels corresponding to the plurality of first objects. The first computing system can cause the graph neural network to generate a plurality of second labels of a first example video and update the graph neural network based at least on the plurality of second labels and the first example video. The first computing system can cause the graph neural network to generate a plurality of third labels of a second example video. The first computing system can output a request for a modification to the at least one third label responsive to the uncertainty score satisfying an annotation criterion..
Any inquiry concerning this communication or earlier communications from the examiner should be directed to XIN SHENG whose telephone number is (571)272-5734. The examiner can normally be reached M-F 9:30AM-3:30PM 6:00PM-8:30PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jason Chan can be reached at 5712723022. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Xin Sheng/Primary Examiner, Art Unit 2619