DETAILED ACTION
Notice of Pre-AIA or AIA Status
Claims 1-20 are pending in this application. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
Examiner Note - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Examiner has reviewed the claims taking into consideration whether they are directed towards statutory subject matter. In this particular application, Claim 19 is directed to a method claim that recites the following limitations:
19. A method, comprising:
accessing a plurality of medical images; and
matching, by a processor, one or more anatomical structures across the plurality of medical images by input of the plurality of medical images to a model trained for patch-matching that captures both global and local patterns embedded within medical images.
The claims were considered in accordance with the 35 U.S.C. 101 analysis and flowchart, and it was determined that the claim constitutes statutory subject matter in accordance with Step 1 as it is directed towards a method. The claims were also considered in accordance with Step 2A, and although on a high level, it does appear to be directed towards an abstract idea and would be a judicial exception, when additional consideration is given to Step 2B, the recitation of the limitations for “a model trained for patch-matching that captures both global and local patterns embedded within medical images” appears to constitute a specialized computing system that incorporates a trained model constituting significantly more than the judicial exception. As a result, the claim was determined to qualify as eligible subject matter and was examined in accordance with the merits.
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-5 and 19-20 are rejected under 35 U.S.C. 102(a)(1)/(a)(2) as being anticipated by Wang et al. (US PGPub US 20240303506, filed March 7, 2024, with provisional prority to March 10, 2023), hereby referred to as “Wang”.
Consider Claims 1, 19 and 20.
Wang teaches:
1. A system for medical image analysis, comprising: / 19. A method, comprising: / 20. A non-transitory, computer-readable medium storing instructions encoded thereon, (Wang: abstract, One embodiment of the present invention sets forth a technique for training a machine learning model to perform feature extraction. The technique includes executing a student version of the machine learning model to generate a first set of features from a first set of image crops and executing a teacher version of the machine learning model to generate a second set of features from a second set of image crops. The technique also includes training the student version of the machine learning model based on one or more losses computed between the first and second sets of features. The technique further includes transmitting the trained student version of the machine learning model to a server, wherein the trained student version can be aggregated by the server with additional trained student versions of the machine learning model to generate a global version of the machine learning model. [0011] Consistent with some embodiments, a computer-implemented method for generating predictions associated with a surgical video includes receiving a global version of a machine learning model, wherein the global version of the machine learning model was generated based on an aggregation of a plurality of local versions of the machine learning model; inputting a sequence of patches extracted from a surgical video into the global version of the machine learning model; executing the global version of the machine learning model to generate predictions of one or more surgical phases or one or more surgical tasks associated with the sequence of patches; and causing the surgical video to be outputted in association with the predictions. [0088], Figures 3A-3B, [0097] FIG. 4 is a flow diagram of method steps for coordinating self-supervised training of a machine learning model at a set of clients, according to various embodiments. Although the method steps are described with respect to the systems of FIGS. 1-2 , persons skilled in the art will understand that any system configured to perform the method steps, in any order, falls within the scope of the various embodiments.)
1. a processor configured to execute one or more operations; and a memory in operable communication with the processor storing instructions the processor executes to execute the one or more operations, to: / 20. the instructions, when executed by one or more processors, cause the one or more processors to perform operations to: (Wang: [0052]-[0064], Figure 2; [0053] Each client 202 includes a processor 242, one or more input/output (I/O) devices 244, and a memory 246, coupled together. Processor 242 includes any technically feasible set of hardware units configured to process data and execute software applications. For example, processor 242 could include one or more CPUs, GPUs, FPGAS, ASICs, DSPs, TPUs, and/or other types of integrated circuits that can be used to process instructions. I/O devices 244 include any technically feasible set of devices configured to perform input and/or output operations, including (but not limited to) a display device, keyboard, touchscreen, mouse, microphone, and/or speaker. [0054] Memory 246 includes any technically feasible storage media configured to store data and software applications. For example, memory 246 could include (but is not limited to) a hard disk, a RAM module, and a ROM module. Memory 246 includes a database 238, a model update engine 212, and an execution engine 252. Database 230 includes a relational database, graph database, key-value store, file system, data warehouse, data storage application, cloud storage, collection of files, and/or another type of data store. Model update engine 212 includes a software application that, when executed by processor 242, interoperates with a corresponding model update engine 210 executing on server 204 to train and/or execute one or more machine learning models.)
1. access a plurality of medical images; / 19. accessing a plurality of medical images; /20. access a plurality of medical images; (Wang: [0031]-[0032], Figure 1; [0032] A display unit 112 is also included in workstation 102. Display unit 112 can display images for viewing by operator 108. Display unit 112 can be moved in various degrees of freedom to accommodate the viewing position of operator 108 and/or to optionally provide control functions as another leader input device. In the example of computer-assisted system 100, displayed images can depict a worksite at which operator 108 is performing various tasks by manipulating leader input devices 106 and/or display unit 112. In some examples, images displayed by display unit 112 can be received by workstation 102 from one or more imaging devices arranged at a worksite. In other examples, the images displayed by display unit 112 can be generated by display unit 112 (or by a different connected device or system), such as for virtual representations of tools, the worksite, or for user interface components.)
1. and match one or more anatomical structures across the plurality of medical images by input of the plurality of medical images to a model trained for patch-matching that captures both global and local patterns embedded within medical images./ 19. and matching, by a processor, one or more anatomical structures across the plurality of medical images by input of the plurality of medical images to a model trained for patch-matching that captures both global and local patterns embedded within medical images./ 20. and match one or more anatomical structures across the plurality of medical images by input of the plurality of medical images to a model trained for patch-matching that captures both global and local patterns embedded within medical images. (Wang: [0011] Consistent with some embodiments, a computer-implemented method for generating predictions associated with a surgical video includes receiving a global version of a machine learning model, wherein the global version of the machine learning model was generated based on an aggregation of a plurality of local versions of the machine learning model; inputting a sequence of patches extracted from a surgical video into the global version of the machine learning model; executing the global version of the machine learning model to generate predictions of one or more surgical phases or one or more surgical tasks associated with the sequence of patches; and causing the surgical video to be outputted in association with the predictions. [0062] Model update engine 210 receives parameters for different local models from clients 202 and stores each set of parameters with an identifier for the corresponding local model in database 230. After a certain number of local models have been received, a certain amount of time has passed, and/or another condition specified in aggregation policy 206 is met, model update engine 210 performs one or more aggregated updates 224 that aggregate parameters from the local models into a new global model 222 (e.g., according to one or more parameters specified in aggregation policy 206). Model update engine 210 also transmits the new global model 222 to clients 202. After model update engine 212 on a given client 202 receives a new global model 222 from server 204, that model update engine 212 trains a new local model using parameters from the new global model 222 as a starting point. The process repeats until a certain number of “global synchronization rounds” involving the aggregation of a set of local models from clients 202 into an updated global model 222 at server 204 have been performed, parameters of global model 222 converge, one or more losses associated with global model 222 fall below a threshold, and/or another condition specified in aggregation policy 206 is met. [0076]-[0083], [0077] As described above, the first set of inputs can include local views of images or video frames in the samples, and the second set of inputs can include both local and global views of the same images or video frames. For example, the first set of inputs ûj s could include multiple sequences of patches extracted from randomized crops of video, where each randomized crop includes less than 50% of the pixels in a corresponding sequence of video frames. The second set of inputs ûj t could include the same sequences of patches as the first set of inputs, as well as additional sequences of patches extracted from larger randomized crops of the same video, where each larger randomized crop includes more than 50% of the pixels in a corresponding sequence of video frames. Each sequence of patches could be denoted by a vector
PNG
media_image1.png
42
109
media_image1.png
Greyscale
, where each patch has dimensions P×P, p denotes a spatial location of a patch within a given crop of a frame, and t represents an index over frames within a sequence of video. [0078] In some embodiments, the ClientUpdate function performs additional augmentations of samples uj from training data 216 to generate inputs ûj s and/or ûj t. For example, the ClientUpdate function could generate one or both sets of inputs ûj s and/or ûj t by applying color perturbations, rotations, CutMix augmentations, and/or other types of changes to samples uj from training data 216. [0079] The ClientUpdate function also computes a consistency loss between a first set of features generated by student model 218 from the first set of inputs and a second set of features generated by teacher model 220 from the second set of inputs.)
Consider Claim 2. Wang teaches: 2. The system of claim 1, wherein the model associated with a self-supervised learning (SSL) framework that defines a Student-Teacher model to extract features of two crops simultaneously. (Wang: [0008]-[0009], [0062] Model update engine 210 receives parameters for different local models from clients 202 and stores each set of parameters with an identifier for the corresponding local model in database 230. After a certain number of local models have been received, a certain amount of time has passed, and/or another condition specified in aggregation policy 206 is met, model update engine 210 performs one or more aggregated updates 224 that aggregate parameters from the local models into a new global model 222 (e.g., according to one or more parameters specified in aggregation policy 206). Model update engine 210 also transmits the new global model 222 to clients 202. After model update engine 212 on a given client 202 receives a new global model 222 from server 204, that model update engine 212 trains a new local model using parameters from the new global model 222 as a starting point. The process repeats until a certain number of “global synchronization rounds” involving the aggregation of a set of local models from clients 202 into an updated global model 222 at server 204 have been performed, parameters of global model 222 converge, one or more losses associated with global model 222 fall below a threshold, and/or another condition specified in aggregation policy 206 is met. [0064] )
Consider Claim 3. Wang teaches: 3. The system of claim 2, wherein the SSL framework includes an image augmentation and restoration module that aims to restore image crops from the two augmentation ways shuffle patches and add noise. (Wang: [0064] Each model update engine 212 can also generate training data 216 that includes different representations of data collected or generated by the corresponding client 202. For example, model update engine 212 at each client 202 could generate X “global” views and Y “local” views of videos of surgical procedures performed at that client 202. [0066] In general, model update engine 212 can generate multiple types of augmentations to training data 216. Each type of augmentation can include a crop that encompasses more or less than a certain proportion of pixels within each frame of the video. Each type of augmentation can also, or instead, include rotations, scalings, shearings, histogram shifts, additions of noise, and/or other transformations of pixels within video and/or crops of the video. Model update engine 212 could use student model 218 to convert a first set of augmentations of training data 216 into corresponding feature maps. Model update engine 212 could also use teacher model 220 to convert a second set of augmentations of training data 216 into corresponding feature maps. The first set of augmentations and the second set of augmentations can include one or more of the same augmentations. The first set of augmentations (or the second set of augmentations) can also, or instead, include one or more augmentations that are not found in the second set of augmentations (or the first set of augmentations).)
Consider Claim 4. Wang teaches: 4. The system of claim 2, wherein the SSL framework includes a global module that aims to enforce the model to learn coarse-grained global features of two crops. (Wang: [0083] For example, training data 214 could be stored in database 230 on server 204 and include segments of videos of surgical procedures that are mapped to labels that represent surgical phases. After training of global model 222 is complete, model update engine 210 could add a classification head that includes a softmax layer and/or other types of neural network layers to global model 222. Model update engine 210 could also perform one or more rounds of supervised fine-tuning of global model 222 using training data 214. More specifically, model update engine 210 could input various segments of videos from training data 214 into global model 222 and obtain predictions of surgical phases, surgical tasks, and/or other types of classes as corresponding output of the classification head added to global model 222. Model update engine 210 could compute a cross-entropy loss, Kullback-Leibler (KL) divergence, and/or another type of classification loss between the output and the corresponding labels. Model update engine 210 could additionally use a training technique (e.g., gradient descent and backpropagation) to update parameters of global model 222, including the newly added classification head, in a way that reduces the computed losses. [0085] After training of a given global model 222 is complete, that global model 222 can be used to process additional data at each client 202 and/or at other locations at which the same types of data are generated or stored. More specifically, model update engine 210 and/or server 204 can store different trained versions of global model 222 in database 230. When a new trained version of global model 222 is available (e.g., after federated self-supervised training and supervised fine-tuning of that version is complete), model update engine 210 and/or server 204 can transmit that version to clients 202 and/or the other locations. An instance of execution engine 252 at each location that receives the trained version of global model 222 can execute the new trained version of global model 222 on a real-time, near-real-time, and/or offline basis to generate predictions 248 of classes and/or other attributes associated with data at that location.)
Consider Claim 5. Wang teaches: 5. The system of claim 2, wherein the SSL framework includes a local module that aims to enforce the model to learn fine-grained local features from overlapped patches. (Wang: [0083] For example, training data 214 could be stored in database 230 on server 204 and include segments of videos of surgical procedures that are mapped to labels that represent surgical phases. After training of global model 222 is complete, model update engine 210 could add a classification head that includes a softmax layer and/or other types of neural network layers to global model 222. Model update engine 210 could also perform one or more rounds of supervised fine-tuning of global model 222 using training data 214. More specifically, model update engine 210 could input various segments of videos from training data 214 into global model 222 and obtain predictions of surgical phases, surgical tasks, and/or other types of classes as corresponding output of the classification head added to global model 222. Model update engine 210 could compute a cross-entropy loss, Kullback-Leibler (KL) divergence, and/or another type of classification loss between the output and the corresponding labels. Model update engine 210 could additionally use a training technique (e.g., gradient descent and backpropagation) to update parameters of global model 222, including the newly added classification head, in a way that reduces the computed losses. [0085] After training of a given global model 222 is complete, that global model 222 can be used to process additional data at each client 202 and/or at other locations at which the same types of data are generated or stored. More specifically, model update engine 210 and/or server 204 can store different trained versions of global model 222 in database 230. When a new trained version of global model 222 is available (e.g., after federated self-supervised training and supervised fine-tuning of that version is complete), model update engine 210 and/or server 204 can transmit that version to clients 202 and/or the other locations. An instance of execution engine 252 at each location that receives the trained version of global model 222 can execute the new trained version of global model 222 on a real-time, near-real-time, and/or offline basis to generate predictions 248 of classes and/or other attributes associated with data at that location.)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made.
Claims 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (US PGPub US 20240303506, filed March 7, 2024, with provisional prority to March 10, 2023), hereby referred to as “Wang”, in view of Miri et al. (US PGPub US 20250299506 A1 filed on June 5, 2025 as a continuation of PCT/US2023/083003 with provisional priority dating to December 8, 2022), hereby referred to as “Miri”.
Consider Claims 1, 19 and 20.
Wang teaches:
1. A system for medical image analysis, comprising: / 19. A method, comprising: / 20. A non-transitory, computer-readable medium storing instructions encoded thereon, (Wang: abstract, One embodiment of the present invention sets forth a technique for training a machine learning model to perform feature extraction. The technique includes executing a student version of the machine learning model to generate a first set of features from a first set of image crops and executing a teacher version of the machine learning model to generate a second set of features from a second set of image crops. The technique also includes training the student version of the machine learning model based on one or more losses computed between the first and second sets of features. The technique further includes transmitting the trained student version of the machine learning model to a server, wherein the trained student version can be aggregated by the server with additional trained student versions of the machine learning model to generate a global version of the machine learning model. [0011] Consistent with some embodiments, a computer-implemented method for generating predictions associated with a surgical video includes receiving a global version of a machine learning model, wherein the global version of the machine learning model was generated based on an aggregation of a plurality of local versions of the machine learning model; inputting a sequence of patches extracted from a surgical video into the global version of the machine learning model; executing the global version of the machine learning model to generate predictions of one or more surgical phases or one or more surgical tasks associated with the sequence of patches; and causing the surgical video to be outputted in association with the predictions. [0088], Figures 3A-3B, [0097] FIG. 4 is a flow diagram of method steps for coordinating self-supervised training of a machine learning model at a set of clients, according to various embodiments. Although the method steps are described with respect to the systems of FIGS. 1-2 , persons skilled in the art will understand that any system configured to perform the method steps, in any order, falls within the scope of the various embodiments.)
1. a processor configured to execute one or more operations; and a memory in operable communication with the processor storing instructions the processor executes to execute the one or more operations, to: / 20. the instructions, when executed by one or more processors, cause the one or more processors to perform operations to: (Wang: [0052]-[0064], Figure 2; [0053] Each client 202 includes a processor 242, one or more input/output (I/O) devices 244, and a memory 246, coupled together. Processor 242 includes any technically feasible set of hardware units configured to process data and execute software applications. For example, processor 242 could include one or more CPUs, GPUs, FPGAS, ASICs, DSPs, TPUs, and/or other types of integrated circuits that can be used to process instructions. I/O devices 244 include any technically feasible set of devices configured to perform input and/or output operations, including (but not limited to) a display device, keyboard, touchscreen, mouse, microphone, and/or speaker. [0054] Memory 246 includes any technically feasible storage media configured to store data and software applications. For example, memory 246 could include (but is not limited to) a hard disk, a RAM module, and a ROM module. Memory 246 includes a database 238, a model update engine 212, and an execution engine 252. Database 230 includes a relational database, graph database, key-value store, file system, data warehouse, data storage application, cloud storage, collection of files, and/or another type of data store. Model update engine 212 includes a software application that, when executed by processor 242, interoperates with a corresponding model update engine 210 executing on server 204 to train and/or execute one or more machine learning models.)
1. access a plurality of medical images; / 19. accessing a plurality of medical images; /20. access a plurality of medical images; (Wang: [0031]-[0032], Figure 1; [0032] A display unit 112 is also included in workstation 102. Display unit 112 can display images for viewing by operator 108. Display unit 112 can be moved in various degrees of freedom to accommodate the viewing position of operator 108 and/or to optionally provide control functions as another leader input device. In the example of computer-assisted system 100, displayed images can depict a worksite at which operator 108 is performing various tasks by manipulating leader input devices 106 and/or display unit 112. In some examples, images displayed by display unit 112 can be received by workstation 102 from one or more imaging devices arranged at a worksite. In other examples, the images displayed by display unit 112 can be generated by display unit 112 (or by a different connected device or system), such as for virtual representations of tools, the worksite, or for user interface components.)
1. and match one or more anatomical structures across the plurality of medical images by input of the plurality of medical images to a model trained for patch-matching that captures both global and local patterns embedded within medical images./ 19. and matching, by a processor, one or more anatomical structures across the plurality of medical images by input of the plurality of medical images to a model trained for patch-matching that captures both global and local patterns embedded within medical images./ 20. and match one or more anatomical structures across the plurality of medical images by input of the plurality of medical images to a model trained for patch-matching that captures both global and local patterns embedded within medical images. (Wang: [0011] Consistent with some embodiments, a computer-implemented method for generating predictions associated with a surgical video includes receiving a global version of a machine learning model, wherein the global version of the machine learning model was generated based on an aggregation of a plurality of local versions of the machine learning model; inputting a sequence of patches extracted from a surgical video into the global version of the machine learning model; executing the global version of the machine learning model to generate predictions of one or more surgical phases or one or more surgical tasks associated with the sequence of patches; and causing the surgical video to be outputted in association with the predictions. [0062] Model update engine 210 receives parameters for different local models from clients 202 and stores each set of parameters with an identifier for the corresponding local model in database 230. After a certain number of local models have been received, a certain amount of time has passed, and/or another condition specified in aggregation policy 206 is met, model update engine 210 performs one or more aggregated updates 224 that aggregate parameters from the local models into a new global model 222 (e.g., according to one or more parameters specified in aggregation policy 206). Model update engine 210 also transmits the new global model 222 to clients 202. After model update engine 212 on a given client 202 receives a new global model 222 from server 204, that model update engine 212 trains a new local model using parameters from the new global model 222 as a starting point. The process repeats until a certain number of “global synchronization rounds” involving the aggregation of a set of local models from clients 202 into an updated global model 222 at server 204 have been performed, parameters of global model 222 converge, one or more losses associated with global model 222 fall below a threshold, and/or another condition specified in aggregation policy 206 is met. [0076]-[0083], [0077] As described above, the first set of inputs can include local views of images or video frames in the samples, and the second set of inputs can include both local and global views of the same images or video frames. For example, the first set of inputs ûj s could include multiple sequences of patches extracted from randomized crops of video, where each randomized crop includes less than 50% of the pixels in a corresponding sequence of video frames. The second set of inputs ûj t could include the same sequences of patches as the first set of inputs, as well as additional sequences of patches extracted from larger randomized crops of the same video, where each larger randomized crop includes more than 50% of the pixels in a corresponding sequence of video frames. Each sequence of patches could be denoted by a vector
PNG
media_image1.png
42
109
media_image1.png
Greyscale
, where each patch has dimensions P×P, p denotes a spatial location of a patch within a given crop of a frame, and t represents an index over frames within a sequence of video. [0078] In some embodiments, the ClientUpdate function performs additional augmentations of samples uj from training data 216 to generate inputs ûj s and/or ûj t. For example, the ClientUpdate function could generate one or both sets of inputs ûj s and/or ûj t by applying color perturbations, rotations, CutMix augmentations, and/or other types of changes to samples uj from training data 216. [0079] The ClientUpdate function also computes a consistency loss between a first set of features generated by student model 218 from the first set of inputs and a second set of features generated by teacher model 220 from the second set of inputs.)
Even if Wang does not specifically teach
4. enforce the model to learn coarse-grained global features of two crops
5. enforce the model to learn fine-grained local features from overlapped patches.
Miri teaches:
1. A system for medical image analysis, comprising: / 19. A method, comprising: / 20. A non-transitory, computer-readable medium storing instructions encoded thereon, (Miri: abstract, The system and method for processing a digital pathology image using a machine learning model that includes a self-supervised hierarchical Vision Transformer (ViT) configured to perform unsupervised clustering with multiple classification tokens. The method includes receiving a digital pathology image that depicts a tissue slice stained with histological dyes. The digital pathology image may be processed to generate a result comprising multiple predicted classifications of individual patches of the digital pathology image. The result is generated by a machine-learning model using a self-supervised hierarchical Vision Transformer (ViT) that may further comprise a multi-head self-attention module configured to predict a crosspatch relevance metric using an attention mechanism for each individual patch in the digital pathology image thereby assigning the individual patches to a cluster based on the crosspatch relevance metrics.)
1. a processor configured to execute one or more operations; and a memory in operable communication with the processor storing instructions the processor executes to execute the one or more operations, to: / 20. the instructions, when executed by one or more processors, cause the one or more processors to perform operations to: (Miri: [0008]-[0009], [0037] FIG. 1 illustrates an example workflow 100 of an SSL model for a given dataset 105. SSL heavily relies on unlabeled data 110 such that instead of having explicit annotations 112, a machine-learning model 120 is trained to create its own understanding of the data by generating auxiliary tasks also referred as pretext tasks 130 that are inherently related to data itself. [0087] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods and/or part or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform part or all of one or more methods and/or part or all of one or more processes disclosed herein.)
1. access a plurality of medical images; / 19. accessing a plurality of medical images; /20. access a plurality of medical images; (Miri: [0037]-[0038], [0039] Hence, in pre-training the model 120 is empowered to extract coarse and/or fine-grained features from an image dataset e.g., digital pathology images. Once the downstream model 140 is also trained, an image can be tested by first giving it to the pre-trained model 120 to extract features and then feeding into the downstream model 140 for downstream tasks (e.g., classification, captioning, segmentation etc.). To reformulate, the key to success of an SSL model may lie in wisely making use of the information derived from the image itself during pre-training. )
1. and match one or more anatomical structures across the plurality of medical images by input of the plurality of medical images to a model trained for patch-matching that captures both global and local patterns embedded within medical images./ 19. and matching, by a processor, one or more anatomical structures across the plurality of medical images by input of the plurality of medical images to a model trained for patch-matching that captures both global and local patterns embedded within medical images./ 20. and match one or more anatomical structures across the plurality of medical images by input of the plurality of medical images to a model trained for patch-matching that captures both global and local patterns embedded within medical images. (Miri: [0038] After pre-training on the SSL pretext tasks 130, the pre-trained model 120 can be fine-tuned on a smaller labeled dataset 115 for specific downstream tasks 150. This transfer learning 135 process leverages the knowledge gained during the self-supervised pre-training to improve performance on downstream tasks 150 that have limited labeled data 115. It is worth mentioning that the pre-trained machine-learning model 120 and the downstream model 140 to be utilized for downstream tasks 150 can be similar or different depending on the specific implementation and requirements. In some instances, the pre-trained model 120 can be used directly for downstream tasks 150. The idea is that the features or representations 125 learned during pre-training can also be useful to perform similar tasks. In other instances, the pre-trained model 120 can be fine-tuned for downstream tasks. Fine-tuning may involve updating the parameters of the pre-trained model 120 adapting to the specific labeled data 115 and downstream tasks 150. Alternatively, a different model can be trained for downstream tasks 150, for example, using a ViT as the pretext model 120 and a convolutional neural network (CNN) as the downstream model 140 for image classification. In this example, ViT is being used as a feature extractor and the CNN as a classifier. [0039] Hence, in pre-training the model 120 is empowered to extract coarse and/or fine-grained features from an image dataset e.g., digital pathology images. Once the downstream model 140 is also trained, an image can be tested by first giving it to the pre-trained model 120 to extract features and then feeding into the downstream model 140 for downstream tasks (e.g., classification, captioning, segmentation etc.). To reformulate, the key to success of an SSL model may lie in wisely making use of the information derived from the image itself during pre-training. [0040] In FIG. 2A, an illustrative example of a self-supervised contrastive method, DINO (Distillation of Information via Non-parametric contrasting) is shown.)
2. The system of claim 1, wherein the model associated with a self-supervised learning (SSL) framework that defines a Student-Teacher model to extract features of two crops simultaneously. (Miri: [0040] In FIG. 2A, an illustrative example of a self-supervised contrastive method, DINO (Distillation of Information via Non-parametric contrasting) is shown. The example architecture of DINO may comprise a student network 210 and a teacher network 215. These two networks student 210 and teacher 215 networks may have similar architecture but different learnable weights due to different update methods. The network 200 learns through a process called knowledge distillation in a self-supervised setting. Distillation may refer to the process of transferring knowledge from a teacher network 215 to a student network 210. The teacher-student network involves training a teacher network 215 to produce reference representations (e.g., z′2, z2) for the training samples 110. The student network 210, in turn, is trained to mimic the representations, and the process can be termed as knowledge distillation. The teacher 215 may be a momentum teacher, which means that the weights of the teacher network 215 are an exponentially weighted average of the weights of student network 210. DINO may use a learning objective 270 for distinguishing the representations of different augmentations of the same image using a memory bank of features from previous instances in the training data. [0041] DINO may define a pretext task that the model needs to learn during training. The pretext task may involve augmenting the input unlabeled data 110 and training the model to distinguish between different augmentations (e.g., V 240 and V′ 245) of the input image in a self- supervised manner. For example, as depicted in FIG. 2 , DINO takes an image x from the unlabeled dataset 110 and apply two different transformations or augmentations 240 and 245 to produce two different views V and V′ that are to be fed into student 210 and teacher network 215 pipelines.)
4. The system of claim 2, wherein the SSL framework includes a global module that aims to enforce the model to learn coarse-grained global features of two crops. (Miri: [0042] A multi-crop augmentation may be applied to extract two sets of images (that may be partially overlapping) from the transformed views V and V′. Small crops may be called local views 220 (e.g., <50% of the image) and large crops (e.g., >50% of the image) may be called global views 225. In other words, the set of global views 225 are of higher dimensions than the set of local views 220. All crops are passed through the student 210 while only the global views 225 are passed through the teacher 215. This encourages “local-to-global” correspondence, training the student 210 to interpolate context from a small crop. During training, only the student 210 is trained so that the set of networks becomes able to understand that the local and global representation, although apparently different, signify the same subject. It is worth mentioning that the multi-crop augmentation and random transformations may be applied in any sequence. For example, local 220 and global views 225 can be achieved from an input image x followed by applying random augmentations (e.g., color jittering, Gaussian blur, solarization etc.) on the local 220 and global views 225 to make the network more robust.)
5. The system of claim 2, wherein the SSL framework includes a local module that aims to enforce the model to learn fine-grained local features from overlapped patches. (Miri: [0042] A multi-crop augmentation may be applied to extract two sets of images (that may be partially overlapping) from the transformed views V and V′. Small crops may be called local views 220 (e.g., <50% of the image) and large crops (e.g., >50% of the image) may be called global views 225. In other words, the set of global views 225 are of higher dimensions than the set of local views 220. All crops are passed through the student 210 while only the global views 225 are passed through the teacher 215. This encourages “local-to-global” correspondence, training the student 210 to interpolate context from a small crop. During training, only the student 210 is trained so that the set of networks becomes able to understand that the local and global representation, although apparently different, signify the same subject. It is worth mentioning that the multi-crop augmentation and random transformations may be applied in any sequence. For example, local 220 and global views 225 can be achieved from an input image x followed by applying random augmentations (e.g., color jittering, Gaussian blur, solarization etc.) on the local 220 and global views 225 to make the network more robust. [0043] Before feeding these views into the vision transformers (ViTs) 235 and 237, the views may be passed into patching and embedding block 230 to get the augmented embedding vectors. Patching and embedding block 230 may convert an image into equal sized patch tokens and perform a set of operations to acquire the corresponding embedding vectors for each patch. These augmented input embedding vectors to ViTs 235 and 237 may represent a sequence of embeddings of patch tokens, a learnable multi-class tokens prepended to the sequence, and the positional information.)
It would have been obvious before the effective filing date of the claimed invention to one of ordinary skill in the art to modify the self-supervised learning model of Wang for federated feature extraction to further leverage Miri’s improvements for a self-supervised hierarchical vision transformer to perform unsupervised clustering with multiple class tokens. The determination of obviousness is predicated upon the following findings: Both references are directed towards the use of self-supervised learning models in medical diagnostic and image analytics, and the references lend themselves for combination. One skilled in the art would have been motivated to modify Wang in order to leverage an improved machine learning model of Miri that uses a self-supervised hierarchical vision transformer configured to predict a cross-patch relevance metric using patch-based image clustering. Furthermore, the prior art collectively includes each element claimed (though not all in the same reference), and one of ordinary skill in the art could have combined the elements in the manner explained above using known engineering design, interface and programming techniques, without changing a “fundamental” operating principle of Wang, while the teaching of Miri continues to perform the same function as originally taught prior to being combined, in order to produce the repeatable and predictable result of improving the self-supervised learning model using coarse-grained global features and fine-grained local features from overlapped patches as suggested by Miri. It is for at least the aforementioned reasons that the examiner has reached a conclusion of obviousness with respect to the claim in question.
Consider Claim 2.
The combination of Wang and Miri teaches:
2. The system of claim 1, wherein the model associated with a self-supervised learning (SSL) framework that defines a Student-Teacher model to extract features of two crops simultaneously. (Wang: [0008]-[0009], [0062] Model update engine 210 receives parameters for different local models from clients 202 and stores each set of parameters with an identifier for the corresponding local model in database 230. After a certain number of local models have been received, a certain amount of time has passed, and/or another condition specified in aggregation policy 206 is met, model update engine 210 performs one or more aggregated updates 224 that aggregate parameters from the local models into a new global model 222 (e.g., according to one or more parameters specified in aggregation policy 206). Model update engine 210 also transmits the new global model 222 to clients 202. After model update engine 212 on a given client 202 receives a new global model 222 from server 204, that model update engine 212 trains a new local model using parameters from the new global model 222 as a starting point. The process repeats until a certain number of “global synchronization rounds” involving the aggregation of a set of local models from clients 202 into an updated global model 222 at server 204 have been performed, parameters of global model 222 converge, one or more losses associated with global model 222 fall below a threshold, and/or another condition specified in aggregation policy 206 is met. [0064] Miri: [0034] For example, a self-supervised multi-class-token hierarchical ViT is a novel backbone that capture both coarse and fine-grained features. (The ViT may have been tested on one or more frameworks, such as the DINO, MOCO, and/or SimCLR SSL frameworks). Compared to ImageNet pretraining and other state-of-the-art SSL methods, this model presents at least two advantages: yielding features of considerably higher quality compared to other state-of-the-art SSL frameworks and tile retrieval demonstration; learning more precise morphological phenotypes down to pixel level, different from the grid structure attention map extracted from the multi-head attention heads from regular ViT. [0035] The unique design of the backbone encoder by expanding the class tokens in regular ViT in a hierarchical manner until the final stage has potential beyond serving as a general-purpose feature extractor in two possible extensions. If trained in SSL paradigm, by equipping each histopathological image with a list of domain-specific attributes as supervisory signals for multiple auxiliary (pretext) tasks (e.g. magnification level, Hematoxylin channel), each class token at the final stage can be customized to predict each label guiding each auxiliary task simultaneously; If trained with some supervisory signals as in weakly supervised settings, each class token at the final stage can be customized to learn targeted lesion regions distinctively.)
Consider Claim 3.
The combination of Wang and Miri teaches:
3. The system of claim 2, wherein the SSL framework includes an image augmentation and restoration module that aims to restore image crops from the two augmentation ways shuffle patches and add noise. (Wang: [0064] Each model update engine 212 can also generate training data 216 that includes different representations of data collected or generated by the corresponding client 202. For example, model update engine 212 at each client 202 could generate X “global” views and Y “local” views of videos of surgical procedures performed at that client 202. [0066] In general, model update engine 212 can generate multiple types of augmentations to training data 216. Each type of augmentation can include a crop that encompasses more or less than a certain proportion of pixels within each frame of the video. Each type of augmentation can also, or instead, include rotations, scalings, shearings, histogram shifts, additions of noise, and/or other transformations of pixels within video and/or crops of the video. Model update engine 212 could use student model 218 to convert a first set of augmentations of training data 216 into corresponding feature maps. Model update engine 212 could also use teacher model 220 to convert a second set of augmentations of training data 216 into corresponding feature maps. The first set of augmentations and the second set of augmentations can include one or more of the same augmentations. The first set of augmentations (or the second set of augmentations) can also, or instead, include one or more augmentations that are not found in the second set of augmentations (or the first set of augmentations). Miri: [0040] The teacher 215 may be a momentum teacher, which means that the weights of the teacher network 215 are an exponentially weighted average of the weights of student network 210. DINO may use a learning objective 270 for distinguishing the representations of different augmentations of the same image using a memory bank of features from previous instances in the training data. [0041] DINO may define a pretext task that the model needs to learn during training. The pretext task may involve augmenting the input unlabeled data 110 and training the model to distinguish between different augmentations (e.g., V 240 and V′ 245) of the input image in a self- supervised manner. For example, as depicted in FIG. 2 , DINO takes an image x from the unlabeled dataset 110 and apply two different transformations or augmentations 240 and 245 to produce two different views V and V′ that are to be fed into student 210 and teacher network 215 pipelines. [0042] A multi-crop augmentation may be applied to extract two sets of images (that may be partially overlapping) from the transformed views V and V′. Small crops may be called local views 220 (e.g., <50% of the image) and large crops (e.g., >50% of the image) may be called global views 225. In other words, the set of global views 225 are of higher dimensions than the set of local views 220. All crops are passed through the student 210 while only the global views 225 are passed through the teacher 215. This encourages “local-to-global” correspondence, training the student 210 to interpolate context from a small crop. During training, only the student 210 is trained so that the set of networks becomes able to understand that the local and global representation, although apparently different, signify the same subject. It is worth mentioning that the multi-crop augmentation and random transformations may be applied in any sequence. For example, local 220 and global views 225 can be achieved from an input image x followed by applying random augmentations (e.g., color jittering, Gaussian blur, solarization etc.) on the local 220 and global views 225 to make the network more robust.)
Consider Claim 4.
The combination of Wang and Miri teaches:
4. The system of claim 2, wherein the SSL framework includes a global module that aims to enforce the model to learn coarse-grained global features of two crops. (Wang: [0083] For example, training data 214 could be stored in database 230 on server 204 and include segments of videos of surgical procedures that are mapped to labels that represent surgical phases. After training of global model 222 is complete, model update engine 210 could add a classification head that includes a softmax layer and/or other types of neural network layers to global model 222. Model update engine 210 could also perform one or more rounds of supervised fine-tuning of global model 222 using training data 214. More specifically, model update engine 210 could input various segments of videos from training data 214 into global model 222 and obtain predictions of surgical phases, surgical tasks, and/or other types of classes as corresponding output of the classification head added to global model 222. Model update engine 210 could compute a cross-entropy loss, Kullback-Leibler (KL) divergence, and/or another type of classification loss between the output and the corresponding labels. Model update engine 210 could additionally use a training technique (e.g., gradient descent and backpropagation) to update parameters of global model 222, including the newly added classification head, in a way that reduces the computed losses. [0085] After training of a given global model 222 is complete, that global model 222 can be used to process additional data at each client 202 and/or at other locations at which the same types of data are generated or stored. More specifically, model update engine 210 and/or server 204 can store different trained versions of global model 222 in database 230. When a new trained version of global model 222 is available (e.g., after federated self-supervised training and supervised fine-tuning of that version is complete), model update engine 210 and/or server 204 can transmit that version to clients 202 and/or the other locations. An instance of execution engine 252 at each location that receives the trained version of global model 222 can execute the new trained version of global model 222 on a real-time, near-real-time, and/or offline basis to generate predictions 248 of classes and/or other attributes associated with data at that location. Miri: [0039] Hence, in pre-training the model 120 is empowered to extract coarse and/or fine-grained features from an image dataset e.g., digital pathology images. Once the downstream model 140 is also trained, an image can be tested by first giving it to the pre-trained model 120 to extract features and then feeding into the downstream model 140 for downstream tasks (e.g., classification, captioning, segmentation etc.). To reformulate, the key to success of an SSL model may lie in wisely making use of the information derived from the image itself during pre-training. [0068] An example implementation of the framework is provided for Cluster-based histopathology Phenotype Representation learning by self-supervised multi-class-token hierarchical ViT (CypherViT) 405 (as illustrated in FIG. 4 ) as the novel backbone encoder to replace the regular ViT 235 in the SSL pipeline inherited from DINO. (It will be appreciated that various other encoders may be used, such as one inherited from MOCO or SimCLR) This approach proves to capture semantically meaningful fine-grained regions of interest detailed to the pixel level. Following the scheme of unsupervised clustering, the single-class token in regular ViT is expanded to a set containing learnable multi-class tokens 330, assembling coarse to fine grained features to semantically aware clusters in a hierarchical manner.)
Consider Claim 5.
The combination of Wang and Miri teaches:
5. The system of claim 2, wherein the SSL framework includes a local module that aims to enforce the model to learn fine-grained local features from overlapped patches. (Wang: [0083] For example, training data 214 could be stored in database 230 on server 204 and include segments of videos of surgical procedures that are mapped to labels that represent surgical phases. After training of global model 222 is complete, model update engine 210 could add a classification head that includes a softmax layer and/or other types of neural network layers to global model 222. Model update engine 210 could also perform one or more rounds of supervised fine-tuning of global model 222 using training data 214. More specifically, model update engine 210 could input various segments of videos from training data 214 into global model 222 and obtain predictions of surgical phases, surgical tasks, and/or other types of classes as corresponding output of the classification head added to global model 222. Model update engine 210 could compute a cross-entropy loss, Kullback-Leibler (KL) divergence, and/or another type of classification loss between the output and the corresponding labels. Model update engine 210 could additionally use a training technique (e.g., gradient descent and backpropagation) to update parameters of global model 222, including the newly added classification head, in a way that reduces the computed losses. [0085] After training of a given global model 222 is complete, that global model 222 can be used to process additional data at each client 202 and/or at other locations at which the same types of data are generated or stored. More specifically, model update engine 210 and/or server 204 can store different trained versions of global model 222 in database 230. When a new trained version of global model 222 is available (e.g., after federated self-supervised training and supervised fine-tuning of that version is complete), model update engine 210 and/or server 204 can transmit that version to clients 202 and/or the other locations. An instance of execution engine 252 at each location that receives the trained version of global model 222 can execute the new trained version of global model 222 on a real-time, near-real-time, and/or offline basis to generate predictions 248 of classes and/or other attributes associated with data at that location. Miri: [0039] Hence, in pre-training the model 120 is empowered to extract coarse and/or fine-grained features from an image dataset e.g., digital pathology images. Once the downstream model 140 is also trained, an image can be tested by first giving it to the pre-trained model 120 to extract features and then feeding into the downstream model 140 for downstream tasks (e.g., classification, captioning, segmentation etc.). To reformulate, the key to success of an SSL model may lie in wisely making use of the information derived from the image itself during pre-training. [0068] An example implementation of the framework is provided for Cluster-based histopathology Phenotype Representation learning by self-supervised multi-class-token hierarchical ViT (CypherViT) 405 (as illustrated in FIG. 4 ) as the novel backbone encoder to replace the regular ViT 235 in the SSL pipeline inherited from DINO. (It will be appreciated that various other encoders may be used, such as one inherited from MOCO or SimCLR) This approach proves to capture semantically meaningful fine-grained regions of interest detailed to the pixel level. Following the scheme of unsupervised clustering, the single-class token in regular ViT is expanded to a set containing learnable multi-class tokens 330, assembling coarse to fine grained features to semantically aware clusters in a hierarchical manner.)
Consider Claim 6.
The combination of Wang and Miri teaches:
6. The system of claim 2, wherein under the SSL framework the model learns coarse-grained, fine-grained and contextualized high-level anatomical structure features. (Wang: [0083] For example, training data 214 could be stored in database 230 on server 204 and include segments of videos of surgical procedures that are mapped to labels that represent surgical phases. After training of global model 222 is complete, model update engine 210 could add a classification head that includes a softmax layer and/or other types of neural network layers to global model 222. Model update engine 210 could also perform one or more rounds of supervised fine-tuning of global model 222 using training data 214. More specifically, model update engine 210 could input various segments of videos from training data 214 into global model 222 and obtain predictions of surgical phases, surgical tasks, and/or other types of classes as corresponding output of the classification head added to global model 222. Model update engine 210 could compute a cross-entropy loss, Kullback-Leibler (KL) divergence, and/or another type of classification loss between the output and the corresponding labels. Model update engine 210 could additionally use a training technique (e.g., gradient descent and backpropagation) to update parameters of global model 222, including the newly added classification head, in a way that reduces the computed losses. [0085] After training of a given global model 222 is complete, that global model 222 can be used to process additional data at each client 202 and/or at other locations at which the same types of data are generated or stored. More specifically, model update engine 210 and/or server 204 can store different trained versions of global model 222 in database 230. When a new trained version of global model 222 is available (e.g., after federated self-supervised training and supervised fine-tuning of that version is complete), model update engine 210 and/or server 204 can transmit that version to clients 202 and/or the other locations. An instance of execution engine 252 at each location that receives the trained version of global model 222 can execute the new trained version of global model 222 on a real-time, near-real-time, and/or offline basis to generate predictions 248 of classes and/or other attributes associated with data at that location. Miri: [0039] Hence, in pre-training the model 120 is empowered to extract coarse and/or fine-grained features from an image dataset e.g., digital pathology images. Once the downstream model 140 is also trained, an image can be tested by first giving it to the pre-trained model 120 to extract features and then feeding into the downstream model 140 for downstream tasks (e.g., classification, captioning, segmentation etc.). To reformulate, the key to success of an SSL model may lie in wisely making use of the information derived from the image itself during pre-training. [0068] An example implementation of the framework is provided for Cluster-based histopathology Phenotype Representation learning by self-supervised multi-class-token hierarchical ViT (CypherViT) 405 (as illustrated in FIG. 4 ) as the novel backbone encoder to replace the regular ViT 235 in the SSL pipeline inherited from DINO. (It will be appreciated that various other encoders may be used, such as one inherited from MOCO or SimCLR) This approach proves to capture semantically meaningful fine-grained regions of interest detailed to the pixel level. Following the scheme of unsupervised clustering, the single-class token in regular ViT is expanded to a set containing learnable multi-class tokens 330, assembling coarse to fine grained features to semantically aware clusters in a hierarchical manner.)
Consider Claim 7.
The combination of Wang and Miri teaches:
7. The system of claim 1, wherein prior to input to the model, the plurality of medical images is pre-processed in grid-wise cropping to get two crops x,x'∈RC×H×W, C is the number of channels, (H,W) are the crops’ spatial dimensions. (Wang: [0064] Each model update engine 212 can also generate training data 216 that includes different representations of data collected or generated by the corresponding client 202. For example, model update engine 212 at each client 202 could generate X “global” views and Y “local” views of videos of surgical procedures performed at that client 202. Each global view could include a crop of a video that includes more than 50% of the pixels within each frame of the video, and each local view could include a crop of a video that includes less than 50% of the pixels within each frame of the video. The crop corresponding to a given global view or local view could include a randomized position, height, and/or width within a given video, subject to any requirements associated with the proportion of pixels in each frame of the video to be included in the crop. A given crop could also, or instead, include a position, height, and/or width that is selected so that the crop captures one or more objects (e.g., surgical instruments, organs, tissue, etc.), activities, and/or other types of detail in the corresponding video and/or portion of a video. [0066] In general, model update engine 212 can generate multiple types of augmentations to training data 216. Each type of augmentation can include a crop that encompasses more or less than a certain proportion of pixels within each frame of the video. Each type of augmentation can also, or instead, include rotations, scalings, shearings, histogram shifts, additions of noise, and/or other transformations of pixels within video and/or crops of the video. Model update engine 212 could use student model 218 to convert a first set of augmentations of training data 216 into corresponding feature maps. Model update engine 212 could also use teacher model 220 to convert a second set of augmentations of training data 216 into corresponding feature maps. The first set of augmentations and the second set of augmentations can include one or more of the same augmentations. The first set of augmentations (or the second set of augmentations) can also, or instead, include one or more augmentations that are not found in the second set of augmentations (or the first set of augmentations).Miri: [0040] The teacher 215 may be a momentum teacher, which means that the weights of the teacher network 215 are an exponentially weighted average of the weights of student network 210. DINO may use a learning objective 270 for distinguishing the representations of different augmentations of the same image using a memory bank of features from previous instances in the training data. [0041] DINO may define a pretext task that the model needs to learn during training. The pretext task may involve augmenting the input unlabeled data 110 and training the model to distinguish between different augmentations (e.g., V 240 and V′ 245) of the input image in a self- supervised manner. For example, as depicted in FIG. 2 , DINO takes an image x from the unlabeled dataset 110 and apply two different transformations or augmentations 240 and 245 to produce two different views V and V′ that are to be fed into student 210 and teacher network 215 pipelines. [0042] A multi-crop augmentation may be applied to extract two sets of images (that may be partially overlapping) from the transformed views V and V′. Small crops may be called local views 220 (e.g., <50% of the image) and large crops (e.g., >50% of the image) may be called global views 225. In other words, the set of global views 225 are of higher dimensions than the set of local views 220. All crops are passed through the student 210 while only the global views 225 are passed through the teacher 215. This encourages “local-to-global” correspondence, training the student 210 to interpolate context from a small crop. During training, only the student 210 is trained so that the set of networks becomes able to understand that the local and global representation, although apparently different, signify the same subject. It is worth mentioning that the multi-crop augmentation and random transformations may be applied in any sequence. For example, local 220 and global views 225 can be achieved from an input image x followed by applying random augmentations (e.g., color jittering, Gaussian blur, solarization etc.) on the local 220 and global views 225 to make the network more robust.
Consider Claim 8.
The combination of Wang and Miri teaches:
8. The system of claim 7, wherein the two crops are input to Student and Teacher encoders fθs,fθt to get the local features s, t respectively. (Miri: [0040] In FIG. 2A, an illustrative example of a self-supervised contrastive method, DINO (Distillation of Information via Non-parametric contrasting) is shown. The example architecture of DINO may comprise a student network 210 and a teacher network 215. These two networks student 210 and teacher 215 networks may have similar architecture but different learnable weights due to different update methods. The network 200 learns through a process called knowledge distillation in a self-supervised setting. Distillation may refer to the process of transferring knowledge from a teacher network 215 to a student network 210. The teacher-student network involves training a teacher network 215 to produce reference representations (e.g., z′2, z2) for the training samples 110. The student network 210, in turn, is trained to mimic the representations, and the process can be termed as knowledge distillation. The teacher 215 may be a momentum teacher, which means that the weights of the teacher network 215 are an exponentially weighted average of the weights of student network 210. DINO may use a learning objective 270 for distinguishing the representations of different augmentations of the same image using a memory bank of features from previous instances in the training data. [0041] DINO may define a pretext task that the model needs to learn during training. The pretext task may involve augmenting the input unlabeled data 110 and training the model to distinguish between different augmentations (e.g., V 240 and V′ 245) of the input image in a self- supervised manner. For example, as depicted in FIG. 2 , DINO takes an image x from the unlabeled dataset 110 and apply two different transformations or augmentations 240 and 245 to produce two different views V and V′ that are to be fed into student 210 and teacher network 215 pipelines. Wang: [0063] As shown in FIG. 2 , each model update engine 212 can train a local model using a student model 218, a teacher model 220, and a set of training data 216 that is generated and/or stored at the corresponding client 202. More specifically, each model update engine 212 can initialize both student model 218 and teacher model 220 to include the parameters and structure of a given version of global model 222 received from server 204. [0064], [0065] Each model update engine 212 can also input different sets of “views” of training data 216 into student model 218 and teacher model 220. Continuing with the above example, model update engine 212 could generate two global views and eight local views of a given video of a surgical procedure. Model update engine 212 could input the global views and local views into teacher model 220 and use teacher model 220 to convert the inputted views into a first set of feature maps. Update engine 212 could input only the local views into student model 218 and use student model 218 to convert the inputted views into a second set of feature maps.)
Consider Claim 9.
The combination of Wang and Miri teaches:
9. The system of claim 8, wherein average pooling operators ⨁:RD×H×W→RD are performed on the local features and the pooled representations are denoted as ys⨁ and yt⨁∈RD. (Wang: [0060] Based on this aggregation policy 206, model update engine 210 initializes a global model 222 for a certain machine learning task and transmits the structure and initial parameters of global model 222 over the corresponding connections to clients 202. Model update engine 212 executing on each client 202 uses global model 222 as a starting point for training a different local model using training data 216 that is generated at that client 202. Additionally, model update engine 212 can use unsupervised training data 216 at the corresponding client 202 to apply unsupervised training updates 228 to the local model at that client 202. [0061] After training of a given local model is complete, the corresponding model update engine 212 transmits parameters of that local model to server 204. Model update engine 212 can also store parameters of the local model with an identifier (e.g., name, version number, timestamp, etc.) for the local model in database 238. [0062] Model update engine 210 receives parameters for different local models from clients 202 and stores each set of parameters with an identifier for the corresponding local model in database 230. After a certain number of local models have been received, a certain amount of time has passed, and/or another condition specified in aggregation policy 206 is met, model update engine 210 performs one or more aggregated updates 224 that aggregate parameters from the local models into a new global model 222 (e.g., according to one or more parameters specified in aggregation policy 206). Model update engine 210 also transmits the new global model 222 to clients 202. After model update engine 212 on a given client 202 receives a new global model 222 from server 204, that model update engine 212 trains a new local model using parameters from the new global model 222 as a starting point. The process repeats until a certain number of “global synchronization rounds” involving the aggregation of a set of local models from clients 202 into an updated global model 222 at server 204 have been performed, parameters of global model 222 converge, one or more losses associated with global model 222 fall below a threshold, and/or another condition specified in aggregation policy 206 is met. [0063] As shown in FIG. 2 , each model update engine 212 can train a local model using a student model 218, a teacher model 220, and a set of training data 216 that is generated and/or stored at the corresponding client 202. More specifically, each model update engine 212 can initialize both student model 218 and teacher model 220 to include the parameters and structure of a given version of global model 222 received from server 204. [0070] An example operation of model update engines 210 and 212 in generating different versions of global model 222, student model 218, and teacher model 220 can be illustrated using the following steps: [0071] PSEUDOCODE In the above pseudocode, model update engines 210 and 212 operate according to input parameters that includes a number of clients 202, a number of global synchronization rounds between model update engine 210 and model update engines 212 on clients 202, and a per-client learning rate. These input parameters can be stored and/or specified in aggregation policy 206 and/or another data source. Miri: [0047] The produced set of projections from MLP is fed to the MLP head Hψ s 260 a for student network and MLP head 260 b for teacher network Hψ t with learnable weights ψ to generate a set of probabilities (q1, q′1) and (q2 and q′2) for the respective student 210 and teacher networks 215. In the context of DINO, the MLP head may represent a component of the projection head, which is responsible for transforming the input features into a space where a learning objective 270 can be applied. The projection heads (260 a and 260 b) may be a layer (e.g., average pooling layer or SoftMax) or a small MLP that takes the embeddings or projections (i.e., 235 d and 237 d) from the respective branches as input and predicts the representations of positive pairs (augmented or transformed views of the same image) and distinguish them from negative pairs (representation from different images). The loss objective 270 aims to maximize the similarity of the two projection sets 235 d and 237 d from the same input while minimizing the similarity to projections of other images within the same mini batch. For contrastive metric measurement, DINO adopts cross entropy loss.)
Consider Claim 10.
The combination of Wang and Miri teaches:
10. The system of claim 1, wherein the processor applying the model in view of the plurality of medical images matches the anatomical structures across different patients. (Miri: [0040] In FIG. 2A, an illustrative example of a self-supervised contrastive method, DINO (Distillation of Information via Non-parametric contrasting) is shown. The example architecture of DINO may comprise a student network 210 and a teacher network 215. These two networks student 210 and teacher 215 networks may have similar architecture but different learnable weights due to different update methods. The network 200 learns through a process called knowledge distillation in a self-supervised setting. Distillation may refer to the process of transferring knowledge from a teacher network 215 to a student network 210. The teacher-student network involves training a teacher network 215 to produce reference representations (e.g., z′2, z2) for the training samples 110. The student network 210, in turn, is trained to mimic the representations, and the process can be termed as knowledge distillation. The teacher 215 may be a momentum teacher, which means that the weights of the teacher network 215 are an exponentially weighted average of the weights of student network 210. DINO may use a learning objective 270 for distinguishing the representations of different augmentations of the same image using a memory bank of features from previous instances in the training data.
[0041] DINO may define a pretext task that the model needs to learn during training. The pretext task may involve augmenting the input unlabeled data 110 and training the model to distinguish between different augmentations (e.g., V 240 and V′ 245) of the input image in a self- supervised manner. For example, as depicted in FIG. 2 , DINO takes an image x from the unlabeled dataset 110 and apply two different transformations or augmentations 240 and 245 to produce two different views V and V′ that are to be fed into student 210 and teacher network 215 pipelines. [0042]…. During training, only the student 210 is trained so that the set of networks becomes able to understand that the local and global representation, although apparently different, signify the same subject. It is worth mentioning that the multi-crop augmentation and random transformations may be applied in any sequence. For example, local 220 and global views 225 can be achieved from an input image x followed by applying random augmentations (e.g., color jittering, Gaussian blur, solarization etc.) on the local 220 and global views 225 to make the network more robust. Wang: [0064] Each model update engine 212 can also generate training data 216 that includes different representations of data collected or generated by the corresponding client 202. For example, model update engine 212 at each client 202 could generate X “global” views and Y “local” views of videos of surgical procedures performed at that client 202. Each global view could include a crop of a video that includes more than 50% of the pixels within each frame of the video, and each local view could include a crop of a video that includes less than 50% of the pixels within each frame of the video. The crop corresponding to a given global view or local view could include a randomized position, height, and/or width within a given video, subject to any requirements associated with the proportion of pixels in each frame of the video to be included in the crop. A given crop could also, or instead, include a position, height, and/or width that is selected so that the crop captures one or more objects (e.g., surgical instruments, organs, tissue, etc.), activities, and/or other types of detail in the corresponding video and/or portion of a video. [0065] Each model update engine 212 can also input different sets of “views” of training data 216 into student model 218 and teacher model 220. Continuing with the above example, model update engine 212 could generate two global views and eight local views of a given video of a surgical procedure. Model update engine 212 could input the global views and local views into teacher model 220 and use teacher model 220 to convert the inputted views into a first set of feature maps. Update engine 212 could input only the local views into student model 218 and use student model 218 to convert the inputted views into a second set of feature maps. [0066]-[0067])
Consider Claim 11.
The combination of Wang and Miri teaches:
11. The system of claim 1, wherein the processor applying the model in view of the plurality of medical images matches the anatomical structures across different views of the same patient. (Miri: [0040] In FIG. 2A, an illustrative example of a self-supervised contrastive method, DINO (Distillation of Information via Non-parametric contrasting) is shown. The example architecture of DINO may comprise a student network 210 and a teacher network 215. These two networks student 210 and teacher 215 networks may have similar architecture but different learnable weights due to different update methods. The network 200 learns through a process called knowledge distillation in a self-supervised setting. Distillation may refer to the process of transferring knowledge from a teacher network 215 to a student network 210. The teacher-student network involves training a teacher network 215 to produce reference representations (e.g., z′2, z2) for the training samples 110. The student network 210, in turn, is trained to mimic the representations, and the process can be termed as knowledge distillation. The teacher 215 may be a momentum teacher, which means that the weights of the teacher network 215 are an exponentially weighted average of the weights of student network 210. DINO may use a learning objective 270 for distinguishing the representations of different augmentations of the same image using a memory bank of features from previous instances in the training data. [0041] DINO may define a pretext task that the model needs to learn during training. The pretext task may involve augmenting the input unlabeled data 110 and training the model to distinguish between different augmentations (e.g., V 240 and V′ 245) of the input image in a self- supervised manner. For example, as depicted in FIG. 2 , DINO takes an image x from the unlabeled dataset 110 and apply two different transformations or augmentations 240 and 245 to produce two different views V and V′ that are to be fed into student 210 and teacher network 215 pipelines. [0042]…. During training, only the student 210 is trained so that the set of networks becomes able to understand that the local and global representation, although apparently different, signify the same subject. It is worth mentioning that the multi-crop augmentation and random transformations may be applied in any sequence. For example, local 220 and global views 225 can be achieved from an input image x followed by applying random augmentations (e.g., color jittering, Gaussian blur, solarization etc.) on the local 220 and global views 225 to make the network more robust. Wang: [0064] Each model update engine 212 can also generate training data 216 that includes different representations of data collected or generated by the corresponding client 202. For example, model update engine 212 at each client 202 could generate X “global” views and Y “local” views of videos of surgical procedures performed at that client 202. Each global view could include a crop of a video that includes more than 50% of the pixels within each frame of the video, and each local view could include a crop of a video that includes less than 50% of the pixels within each frame of the video. The crop corresponding to a given global view or local view could include a randomized position, height, and/or width within a given video, subject to any requirements associated with the proportion of pixels in each frame of the video to be included in the crop. A given crop could also, or instead, include a position, height, and/or width that is selected so that the crop captures one or more objects (e.g., surgical instruments, organs, tissue, etc.), activities, and/or other types of detail in the corresponding video and/or portion of a video. [0065] Each model update engine 212 can also input different sets of “views” of training data 216 into student model 218 and teacher model 220. Continuing with the above example, model update engine 212 could generate two global views and eight local views of a given video of a surgical procedure. Model update engine 212 could input the global views and local views into teacher model 220 and use teacher model 220 to convert the inputted views into a first set of feature maps. Update engine 212 could input only the local views into student model 218 and use student model 218 to convert the inputted views into a second set of feature maps. [0066]-[0067])
Consider Claim 12.
The combination of Wang and Miri teaches:
12. The system of claim 1, wherein the model calculates the consistency loss based on the absolute positions of overlapping image patches of the plurality of medical images. (Wang: [0052] In addition to position and patch embeddings 340, a special token called class [cls] token 330 may be introduced. The semantic image layout can be discovered from the attention maps of the class tokens. These attention maps may lead to promising results in unsupervised segmentation tasks. In some embodiments, unlike regular transformers, multiple class tokens 330 are used. Using a single class token may be challenging for accurate localization of different objects on a single image. Therefore, instead of a single class token multiple class tokens 330 may be used, which will be responsible for learning representations for different object classes. By doing so, the model can learn to attend to the regions of the image that belong to each class and generate class-discriminative object localization maps from the class-to-patch attentions. This technique can be useful for weakly supervised semantic segmentation, which is the task of assigning a class label to each pixel in an image using only image-level labels as supervision. The output of the linear projection block 345, a combination of patch embedding 340, positional encodings 325, and multi-class tokens 330 forms the input to the VIT. [0076]-[0083], [0077] As described above, the first set of inputs can include local views of images or video frames in the samples, and the second set of inputs can include both local and global views of the same images or video frames. For example, the first set of inputs ûj s could include multiple sequences of patches extracted from randomized crops of video, where each randomized crop includes less than 50% of the pixels in a corresponding sequence of video frames. The second set of inputs ûj t could include the same sequences of patches as the first set of inputs, as well as additional sequences of patches extracted from larger randomized crops of the same video, where each larger randomized crop includes more than 50% of the pixels in a corresponding sequence of video frames. Each sequence of patches could be denoted by a vector
PNG
media_image1.png
42
109
media_image1.png
Greyscale
, where each patch has dimensions P×P, p denotes a spatial location of a patch within a given crop of a frame, and t represents an index over frames within a sequence of video. [0078] In some embodiments, the ClientUpdate function performs additional augmentations of samples uj from training data 216 to generate inputs ûj s and/or ûj t. For example, the ClientUpdate function could generate one or both sets of inputs ûj s and/or ûj t by applying color perturbations, rotations, CutMix augmentations, and/or other types of changes to samples uj from training data 216. [0079] The ClientUpdate function also computes a consistency loss between a first set of features generated by student model 218 from the first set of inputs and a second set of features generated by teacher model 220 from the second set of inputs. Miri: [0050] The loss in Equation 1 may be adapted to self-supervised learning problem that deploys a multi-crop strategy with local 220 and global views 225 augmented from the original input image x, as in FIG. 2A. In some embodiments, the multi-crop augmentation is important but with an optimal sweet spot on the number of local views that are treated as tunable hyperparameters. The global views may be denoted as Σgϵ[1,N g ]xg and several local views of smaller resolution are denoted as Σlϵ[1,N 1 ]xg where Nl may refer to the number of local views and Ng may refer to the number of global views. For simplicity, only two global views are demonstrated in FIG. 2A. The loss objective function 270 as given in Equation 1 can be modified as given in Equation 2, where Hψ t and Hψ t , denote the output probability distributions of teacher 215 and student 210 network, respectively. The gradient propagation is stopped in teacher network 215, gradients 265 are only allowed to pass through the student network 210.
PNG
media_image2.png
40
586
media_image2.png
Greyscale
Miri: [0028]
This disclosure describes various elements (such as systems and devices, and portions of systems and devices) in three-dimensional space. As used herein, the term “position” refers to the location of an element or a portion of an element in a three-dimensional space (e.g., three degrees of translational freedom along Cartesian x-, y-, and z-coordinates). As used herein, the term “orientation” refers to the rotational placement of an element or a portion of an element (three degrees of rotational freedom—e.g., roll, pitch, and yaw). As used herein, the term “pose” refers to the multi-degree of freedom (DOF) spatial position and/or orientation of a coordinate system of interest attached to a rigid body. In general, a pose can include a pose variable for each of the DOFs in the pose. For example, a full 6-DOF pose would include 6 pose variables corresponding to the 3 positional DOFs (e.g., x, y, and z) and the 3 orientational DOFs (e.g., roll, pitch, and yaw). A 3-DOF position only pose would include only pose variables for the 3 positional DOFs. Similarly, a 3-DOF orientation only pose would include only pose variables for the 3 rotational DOFs. Poses with any other number of DOFs (e.g., one, two, four, or five) are also possible. As used herein, the term “shape” refers to a set positions or orientations measured along an element. As used herein, and for an element or portion of an element, e.g., a device (e.g., a computer-assisted system or a repositionable arm), the term “proximal” refers to a direction toward the base of the system or device of the repositionable arm along its kinematic chain, and the term “distal” refers to a direction away from the base along the kinematic chain.)
Consider Claim 13.
The combination of Wang and Miri teaches:
13. The system of claim 1, wherein the model as trained: takes, utilizing a student-teacher architecture, a first crop and a second crop from overlapped patches of an image; and learns high-level relationships among anatomical structures by patch order classification and fine-grained image features by patch appearance restoration. (Miri: [0040] In FIG. 2A, an illustrative example of a self-supervised contrastive method, DINO (Distillation of Information via Non-parametric contrasting) is shown. The example architecture of DINO may comprise a student network 210 and a teacher network 215. These two networks student 210 and teacher 215 networks may have similar architecture but different learnable weights due to different update methods. The network 200 learns through a process called knowledge distillation in a self-supervised setting. Distillation may refer to the process of transferring knowledge from a teacher network 215 to a student network 210. The teacher-student network involves training a teacher network 215 to produce reference representations (e.g., z′2, z2) for the training samples 110. The student network 210, in turn, is trained to mimic the representations, and the process can be termed as knowledge distillation. The teacher 215 may be a momentum teacher, which means that the weights of the teacher network 215 are an exponentially weighted average of the weights of student network 210. DINO may use a learning objective 270 for distinguishing the representations of different augmentations of the same image using a memory bank of features from previous instances in the training data. [0041] DINO may define a pretext task that the model needs to learn during training. The pretext task may involve augmenting the input unlabeled data 110 and training the model to distinguish between different augmentations (e.g., V 240 and V′ 245) of the input image in a self- supervised manner. For example, as depicted in FIG. 2 , DINO takes an image x from the unlabeled dataset 110 and apply two different transformations or augmentations 240 and 245 to produce two different views V and V′ that are to be fed into student 210 and teacher network 215 pipelines. [0042]…. During training, only the student 210 is trained so that the set of networks becomes able to understand that the local and global representation, although apparently different, signify the same subject. It is worth mentioning that the multi-crop augmentation and random transformations may be applied in any sequence. For example, local 220 and global views 225 can be achieved from an input image x followed by applying random augmentations (e.g., color jittering, Gaussian blur, solarization etc.) on the local 220 and global views 225 to make the network more robust. Wang: [0064] Each model update engine 212 can also generate training data 216 that includes different representations of data collected or generated by the corresponding client 202. For example, model update engine 212 at each client 202 could generate X “global” views and Y “local” views of videos of surgical procedures performed at that client 202. Each global view could include a crop of a video that includes more than 50% of the pixels within each frame of the video, and each local view could include a crop of a video that includes less than 50% of the pixels within each frame of the video. The crop corresponding to a given global view or local view could include a randomized position, height, and/or width within a given video, subject to any requirements associated with the proportion of pixels in each frame of the video to be included in the crop. A given crop could also, or instead, include a position, height, and/or width that is selected so that the crop captures one or more objects (e.g., surgical instruments, organs, tissue, etc.), activities, and/or other types of detail in the corresponding video and/or portion of a video. [0065] Each model update engine 212 can also input different sets of “views” of training data 216 into student model 218 and teacher model 220. Continuing with the above example, model update engine 212 could generate two global views and eight local views of a given video of a surgical procedure. Model update engine 212 could input the global views and local views into teacher model 220 and use teacher model 220 to convert the inputted views into a first set of feature maps. Update engine 212 could input only the local views into student model 218 and use student model 218 to convert the inputted views into a second set of feature maps. [0066]-[0067])
Consider Claim 14.
The combination of Wang and Miri teaches:
14. The system of claim 13, wherein the model integrates the first crop with the second crop to learn consistent contextualized embedding for coarse-grained global anatomical structures. (Wang: [0083] For example, training data 214 could be stored in database 230 on server 204 and include segments of videos of surgical procedures that are mapped to labels that represent surgical phases. After training of global model 222 is complete, model update engine 210 could add a classification head that includes a softmax layer and/or other types of neural network layers to global model 222. Model update engine 210 could also perform one or more rounds of supervised fine-tuning of global model 222 using training data 214. More specifically, model update engine 210 could input various segments of videos from training data 214 into global model 222 and obtain predictions of surgical phases, surgical tasks, and/or other types of classes as corresponding output of the classification head added to global model 222. Model update engine 210 could compute a cross-entropy loss, Kullback-Leibler (KL) divergence, and/or another type of classification loss between the output and the corresponding labels. Model update engine 210 could additionally use a training technique (e.g., gradient descent and backpropagation) to update parameters of global model 222, including the newly added classification head, in a way that reduces the computed losses. [0085] After training of a given global model 222 is complete, that global model 222 can be used to process additional data at each client 202 and/or at other locations at which the same types of data are generated or stored. More specifically, model update engine 210 and/or server 204 can store different trained versions of global model 222 in database 230. When a new trained version of global model 222 is available (e.g., after federated self-supervised training and supervised fine-tuning of that version is complete), model update engine 210 and/or server 204 can transmit that version to clients 202 and/or the other locations. An instance of execution engine 252 at each location that receives the trained version of global model 222 can execute the new trained version of global model 222 on a real-time, near-real-time, and/or offline basis to generate predictions 248 of classes and/or other attributes associated with data at that location. Miri: [0039] Hence, in pre-training the model 120 is empowered to extract coarse and/or fine-grained features from an image dataset e.g., digital pathology images. Once the downstream model 140 is also trained, an image can be tested by first giving it to the pre-trained model 120 to extract features and then feeding into the downstream model 140 for downstream tasks (e.g., classification, captioning, segmentation etc.). To reformulate, the key to success of an SSL model may lie in wisely making use of the information derived from the image itself during pre-training. [0068] An example implementation of the framework is provided for Cluster-based histopathology Phenotype Representation learning by self-supervised multi-class-token hierarchical ViT (CypherViT) 405 (as illustrated in FIG. 4 ) as the novel backbone encoder to replace the regular ViT 235 in the SSL pipeline inherited from DINO. (It will be appreciated that various other encoders may be used, such as one inherited from MOCO or SimCLR) This approach proves to capture semantically meaningful fine-grained regions of interest detailed to the pixel level. Following the scheme of unsupervised clustering, the single-class token in regular ViT is expanded to a set containing learnable multi-class tokens 330, assembling coarse to fine grained features to semantically aware clusters in a hierarchical manner.)
Consider Claim 15.
The combination of Wang and Miri teaches:
15. The system of claim 1, wherein analogous regions of the plurality of medical images are captured by the first and second crops so that global embedding consistency encourages extraction of features of similar local regions. (Miri: [0038] After pre-training on the SSL pretext tasks 130, the pre-trained model 120 can be fine-tuned on a smaller labeled dataset 115 for specific downstream tasks 150. This transfer learning 135 process leverages the knowledge gained during the self-supervised pre-training to improve performance on downstream tasks 150 that have limited labeled data 115. It is worth mentioning that the pre-trained machine-learning model 120 and the downstream model 140 to be utilized for downstream tasks 150 can be similar or different depending on the specific implementation and requirements. In some instances, the pre-trained model 120 can be used directly for downstream tasks 150. The idea is that the features or representations 125 learned during pre-training can also be useful to perform similar tasks. In other instances, the pre-trained model 120 can be fine-tuned for downstream tasks. Fine-tuning may involve updating the parameters of the pre-trained model 120 adapting to the specific labeled data 115 and downstream tasks 150. Alternatively, a different model can be trained for downstream tasks 150, for example, using a ViT as the pretext model 120 and a convolutional neural network (CNN) as the downstream model 140 for image classification. In this example, ViT is being used as a feature extractor and the CNN as a classifier. [0039] Hence, in pre-training the model 120 is empowered to extract coarse and/or fine-grained features from an image dataset e.g., digital pathology images. Once the downstream model 140 is also trained, an image can be tested by first giving it to the pre-trained model 120 to extract features and then feeding into the downstream model 140 for downstream tasks (e.g., classification, captioning, segmentation etc.). To reformulate, the key to success of an SSL model may lie in wisely making use of the information derived from the image itself during pre-training. [0040] In FIG. 2A, an illustrative example of a self-supervised contrastive method, DINO (Distillation of Information via Non-parametric contrasting) is shown.)
Consider Claim 16.
The combination of Wang and Miri teaches:
16. The system of claim 1, wherein the model learns fine-grained and precise anatomical structures from local patch embeddings of overlapped parts. (Wang: [0083] For example, training data 214 could be stored in database 230 on server 204 and include segments of videos of surgical procedures that are mapped to labels that represent surgical phases. After training of global model 222 is complete, model update engine 210 could add a classification head that includes a softmax layer and/or other types of neural network layers to global model 222. Model update engine 210 could also perform one or more rounds of supervised fine-tuning of global model 222 using training data 214. More specifically, model update engine 210 could input various segments of videos from training data 214 into global model 222 and obtain predictions of surgical phases, surgical tasks, and/or other types of classes as corresponding output of the classification head added to global model 222. Model update engine 210 could compute a cross-entropy loss, Kullback-Leibler (KL) divergence, and/or another type of classification loss between the output and the corresponding labels. Model update engine 210 could additionally use a training technique (e.g., gradient descent and backpropagation) to update parameters of global model 222, including the newly added classification head, in a way that reduces the computed losses. [0085] After training of a given global model 222 is complete, that global model 222 can be used to process additional data at each client 202 and/or at other locations at which the same types of data are generated or stored. More specifically, model update engine 210 and/or server 204 can store different trained versions of global model 222 in database 230. When a new trained version of global model 222 is available (e.g., after federated self-supervised training and supervised fine-tuning of that version is complete), model update engine 210 and/or server 204 can transmit that version to clients 202 and/or the other locations. An instance of execution engine 252 at each location that receives the trained version of global model 222 can execute the new trained version of global model 222 on a real-time, near-real-time, and/or offline basis to generate predictions 248 of classes and/or other attributes associated with data at that location. Miri: [0039] Hence, in pre-training the model 120 is empowered to extract coarse and/or fine-grained features from an image dataset e.g., digital pathology images. Once the downstream model 140 is also trained, an image can be tested by first giving it to the pre-trained model 120 to extract features and then feeding into the downstream model 140 for downstream tasks (e.g., classification, captioning, segmentation etc.). To reformulate, the key to success of an SSL model may lie in wisely making use of the information derived from the image itself during pre-training. [0068] An example implementation of the framework is provided for Cluster-based histopathology Phenotype Representation learning by self-supervised multi-class-token hierarchical ViT (CypherViT) 405 (as illustrated in FIG. 4 ) as the novel backbone encoder to replace the regular ViT 235 in the SSL pipeline inherited from DINO. (It will be appreciated that various other encoders may be used, such as one inherited from MOCO or SimCLR) This approach proves to capture semantically meaningful fine-grained regions of interest detailed to the pixel level. Following the scheme of unsupervised clustering, the single-class token in regular ViT is expanded to a set containing learnable multi-class tokens 330, assembling coarse to fine grained features to semantically aware clusters in a hierarchical manner.)
Consider Claim 17.
The combination of Wang and Miri teaches:
17. The system of claim 1, wherein the model defines a network that that considers both global and local features of medical images at the same time. (Wang: [0060]-[0063], [0060]
To address the above shortcomings, model update engines 210 and 212 include functionality to train one or more machine learning models under a federated self-supervised learning framework. Within the federated self-supervised learning framework, each client 202 establishes a connection over network 250 with server 204. Model update engine 210 executing on server 204 operates according to an aggregation policy 206 associated with the federated self-supervised learning framework. Based on this aggregation policy 206, model update engine 210 initializes a global model 222 for a certain machine learning task and transmits the structure and initial parameters of global model 222 over the corresponding connections to clients 202. Model update engine 212 executing on each client 202 uses global model 222 as a starting point for training a different local model using training data 216 that is generated at that client 202. Additionally, model update engine 212 can use unsupervised training data 216 at the corresponding client 202 to apply unsupervised training updates 228 to the local model at that client 202. [0061] After training of a given local model is complete, the corresponding model update engine 212 transmits parameters of that local model to server 204. Model update engine 212 can also store parameters of the local model with an identifier (e.g., name, version number, timestamp, etc.) for the local model in database 238. [0062] Model update engine 210 receives parameters for different local models from clients 202 and stores each set of parameters with an identifier for the corresponding local model in database 230. After a certain number of local models have been received, a certain amount of time has passed, and/or another condition specified in aggregation policy 206 is met, model update engine 210 performs one or more aggregated updates 224 that aggregate parameters from the local models into a new global model 222 (e.g., according to one or more parameters specified in aggregation policy 206). Model update engine 210 also transmits the new global model 222 to clients 202. After model update engine 212 on a given client 202 receives a new global model 222 from server 204, that model update engine 212 trains a new local model using parameters from the new global model 222 as a starting point. The process repeats until a certain number of “global synchronization rounds” involving the aggregation of a set of local models from clients 202 into an updated global model 222 at server 204 have been performed, parameters of global model 222 converge, one or more losses associated with global model 222 fall below a threshold, and/or another condition specified in aggregation policy 206 is met. Miri: [0042] A multi-crop augmentation may be applied to extract two sets of images (that may be partially overlapping) from the transformed views V and V′. Small crops may be called local views 220 (e.g., <50% of the image) and large crops (e.g., >50% of the image) may be called global views 225. In other words, the set of global views 225 are of higher dimensions than the set of local views 220. All crops are passed through the student 210 while only the global views 225 are passed through the teacher 215. This encourages “local-to-global” correspondence, training the student 210 to interpolate context from a small crop. During training, only the student 210 is trained so that the set of networks becomes able to understand that the local and global representation, although apparently different, signify the same subject. It is worth mentioning that the multi-crop augmentation and random transformations may be applied in any sequence. For example, local 220 and global views 225 can be achieved from an input image x followed by applying random augmentations (e.g., color jittering, Gaussian blur, solarization etc.) on the local 220 and global views 225 to make the network more robust. [0043] Before feeding these views into the vision transformers (ViTs) 235 and 237, the views may be passed into patching and embedding block 230 to get the augmented embedding vectors. Patching and embedding block 230 may convert an image into equal sized patch tokens and perform a set of operations to acquire the corresponding embedding vectors for each patch. These augmented input embedding vectors to ViTs 235 and 237 may represent a sequence of embeddings of patch tokens, a learnable multi-class tokens prepended to the sequence, and the positional information.)
Consider Claim 18.
The combination of Wang and Miri teaches:
18. The system of claim 1, wherein the model localizes arbitrary anatomical structures across views of the same patient and across patients of different genders and weights and of health and disease. (Wang: [0074] The ServerUpdate function then aggregates updates received from clients 202 as a weighted sum of the differences. Each weight w; in the weighted sum is associated with a corresponding client 202 denoted by i. Weights in the weighted sum can be set to equal values, determined based on the number of samples used to train the corresponding student models, and/or based on other factors specified in aggregation policy 206. [0075] The ServerUpdate function also generates a new version of global model 222 by adding the weighted sum of updates from clients 202 to the current set of weights for global model 222. The ServerUpdate function then repeats the process with the new version of global model 222 until the specified number of global synchronization rounds has been performed. [0076] The ClientUpdate function is executed using training data 216 U at the corresponding client 202 and a certain number n of local training iterations. The ClientUpdate function begins by initializing student model 218 θs and teacher model 220 θt using the structure and weights of the most recent global model 222 θ received from server 204. [0097] FIG. 4 is a flow diagram of method steps for coordinating self-supervised training of a machine learning model at a set of clients, according to various embodiments. [0098]As shown, in step 402, model update engine 210 executing on server 204 initializes a global version of a machine learning model. For example, model update engine 210 could initialize neural network layers, blocks, and/or other structures within the machine learning model. Model update engine 210 could also initialize neural network weights of the machine learning model to random values, weights from a pretrained model, and/or weights from a previous version of the machine learning model .Miri: [0045] FIG. 2B shows an illustrative example of a ViTs from FIG. 2A. The two vision transformers (ViTs) 235 and 237 in student and teacher network may have same architecture but different learnable weights due to different update method. A vision transformer (ViT) may be a type of neural network based on transformer architecture. The augmented embedded vectors from patching and embedding 230 are fed into the ViTs that may further include transformer encoders (Ees) 235 a and a multi-layer perceptron (MLP) 235 c in student network 210, a transformer encoder 237 a (Ee,) and an MLP 237 c in teacher network 215 with learnable weights θ. These transformer encoders may represent a stack of multiple self-attention layers. Self-attention may refer to a mechanism that allows the model to learn long-range dependencies between the patches for tasks such as image classification, as it may allow the model to learn how the different parts of an image may contribute to its overall label. The output of the transformer encoder is a sequence of vectors e.g., for student network 210, (y1, y′1) 235 b representing the intermediate features of the set of global 225 and local views 220 and for teacher network 215, (y2, y′2) 237 b representing the intermediate features of the set of global views 225.)
Conclusion
The prior art made of record in form PTO-892 and not relied upon is considered pertinent to applicant's disclosure.
PNG
media_image3.png
238
908
media_image3.png
Greyscale
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TAHMINA ANSARI whose telephone number is 571-270-3379. The examiner can normally be reached on IFP Flex - Monday through Friday 9 to 5.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, O’NEAL MISTRY can be reached on 313-446-4912. The fax phone numbers for the organization where this application or proceeding is assigned are 571-273-8300 for regular communications and 571-273-8300 for After Final communications. TC 2600’s customer service number is 571-272-2600.
Any inquiry of a general nature or relating to the status of this application or proceeding should be directed to the receptionist whose telephone number is 571-272-2600.
2674
/Tahmina Ansari/
August 3, 2026
/TAHMINA N ANSARI/Primary Examiner, Art Unit 2674