DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Preliminary Remarks
This is a reply to the amendments filed on 04/17/2026, in which, claims 1, 6, 13-14, and 19-20 are amended; and claim 5 is cancelled. Claims 1-4 and 6-20 remain pending in the present application with claims 1, 13, and 19 being independent claims.
When making claim amendments, the applicant is encouraged to consider the references in their entireties, including those portions that have not been cited by the examiner and their equivalents as they may most broadly and appropriately apply to any particular anticipated claim amendments.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on February 25, 2026 is in compliance with the provisions of 37 CFR 1.97 and is being considered by the Examiner.
Response to Arguments
Regarding the 35 U.S.C. §112(f) invocation of claims 1, 7, 8, and 9, Applicants have not amended the claims to replace the “device” in claims. The claim limitations still use generic placeholders that are coupled with functional language without reciting sufficient structures to perform the recited functions and the generic placeholders are not preceded by a structural modifier. Therefore, the outstanding 35 U.S.C. §112(f), invocation of claims 1-4 and 6-12 is maintained.
Applicant's arguments filed on 04/17/2026 with respect to amended claims 1, 13, and 19 have been considered but are moot in view of the new ground(s) of rejection.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. - An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
Use of the word “device” (or “step for”, “unit”, “element”, “mechanism”, “module”, “means”, “engine”, “component”, “member”, “apparatus”, “machine”, “system”, “assembly”, “portion”) in a claim with functional language creates a rebuttable presumption that the claim element is to be treated in accordance with 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is invoked is rebutted when the function is recited with sufficient structure, material, or acts within the claim itself to entirely perform the recited function.
Absence of the word “device” in a claim creates a rebuttable presumption that the claim element is not to be treated in accordance with 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is not invoked is rebutted when the claim element recites function but fails to recite sufficiently definite structure, material or acts to perform that function.
The claim limitations use a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are:
a hardware storage device for storing instructions in claim 1; and
an output device for outputting the mixed-modality summary in claim 7.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. (FP 7.30.06).
For more information, see MPEP § 2173 et seq. and Supplementary Examination Guidelines for Determining Compliance With 35 U.S.C. 112 and for Treatment of Related Issues in Patent Applications, 76 FR 7162, 7167 (Feb. 9, 2011).
Claims 2-4 and 6-12 depend on claim 1 thus 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is also invoked.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-2, 6-9, 12-13, and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Jin et al. (US 20230386208 A1, hereinafter referred to as “Jin”) in view of Loo et al. (US 20240212328 A1, hereinafter referred to as “Loo”), and further in view of Crabtree et al. (US 20240386015 A1, hereinafter referred to as “Crabtree”).
Regarding claim 1, Jin discloses a system for generating a mixed-modality summary, comprising:
a hardware storage device for storing instructions that, when executed, cause the system to perform operations (see Jin, paragraph [0030]: “the processor is configured to execute computer-readable instructions stored in a memory to perform various functions”) comprising:
acquiring mixed-modality data covering data from a first number of modalities (see Jin, paragraph [0053]: “the unsupervised temporal segmentation method take input from multiple modalities, including both visual features and language features”); and
outputting the mixed-modality summary (see Jin, paragraph [0073]: “the system can provide the video summary directly to the user through an application, or to a storage such as a database”).
Regarding claim 1, Jin discloses all the claimed limitations with the exception of generating mixed-modality embeddings within a joint embedding space using the mixed-modality data; generating a coreset of the mixed-modality embeddings, the coreset comprises a representative subset of the mixed-modality embeddings; determining a user-derived constraint for an output application; generating a second coreset of embedding vectors using the coreset of the mixed-modality embeddings and the user-derived constraint, each embedding vector of the second coreset of embedding vectors has fewer modalities than the first number of modalities; and generating the mixed-modality summary using the second coreset of embedding vectors.
Loo from the same or similar fields of endeavor discloses generating a coreset of the mixed-modality embeddings, the coreset comprises a representative subset of the mixed-modality embeddings (see Loo, paragraph [0059]: “Once the starting dataset, also referred to as a dataset or an original dataset (though as described herein an original dataset can also be a subset of possible starting datasets), is sampled to form a coreset, a non-deterministic feature neural network training kernel can be applied to at least some portion of the coreset to define a modified coreset”);
each embedding vector of the second coreset of embedding vectors has fewer modalities than the first number of modalities (see Loo, paragraph [0064]: “Synthetic data can be defined as output from the dataset distillation techniques described herein, including datasets that have been created by the output of an algorithm that takes in an input dataset and generates a modified and/or synthetic version of the same dataset. Examples include, but are not limited to, inputting in a dataset and outputting a condensed, distilled version, and/or outputting a private version of the dataset. Therefore, the disclosed techniques receive an input dataset and then creates a version of that specific dataset”); and
generating the mixed-modality summary using the second coreset of embedding vectors (see Loo, paragraph [0085]: “the computer system can initialize a coreset with random images from a dataset (block 302). A batch of images and labels can then be sampled from the dataset (block 304). Any variety of requirements and/or parameters can be used to determine which images are sampled. For example, a relevant user can implement one or more requirements and/or parameters specific to their particular use case”).
Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings as in Loo with the teachings as in Jin. The motivation for doing so would ensure the system to have the ability to use system and method for performing dataset distillation disclosed in Loo to generate a subset of possible starting datasets which is sampled to form a coreset; to generate a condensed, distilled version, and/or outputting a private version of the dataset using distillation techniques wherein the version of that specific dataset has fewer modalities; to determine which images are sampled and to summarize large sets of data into smaller sets while accurately representing the full dataset thus generating a subset of the mixed-modality embeddings wherein each embedding vector of the generated embedding vectors has fewer modalities than the first number of modalities and generating the mixed-modality summary using the subset of embedding vectors in order to generate a mixed-modality summary using the coreset so that mixed-modality summarization can be done more quickly and efficiently.
Regarding claim 1, the combination teachings of Jin and Loo disclose all the claimed limitations with the exception of generating mixed-modality embeddings within a joint embedding space using the mixed-modality data; determining a user-derived constraint for an output application; and generating a second coreset of embedding vectors using the coreset of the mixed-modality embeddings and the user-derived constraint.
Crabtree from the same or similar fields of endeavor discloses generating mixed-modality embeddings within a joint embedding space using the mixed-modality data (see Crabtree, paragraph [0155]: “multi-modal computing system 2123 is present and configured to align and synchronize representations across different data modalities (e.g., text, images, audio, etc.) to create a unified and consistent representation of the input data. Multi-modal system 2123 may implement techniques such as cross-modal attention, multi-modal fusion, and/or joint embedding spaces to effectively combine information from different modalities”);
determining a user-derived constraint for an output application (see Crabtree, paragraph [0387]: “the user-defined rules/preferences may be defined by an entity (e.g., a company). Exemplary rules or preferences can include, but are not limited to, conditional generation preferences, formatting rules, language rules, style rules, geographic rules, environmental rules, and timing rules”); and
generating a second coreset of embedding vectors using the coreset of the mixed-modality embeddings and the user-derived constraint (see Crabtree, paragraphs [0391]-[0392]: “obtained plurality of context data may be processed into vectors by an embedding model and stored in the vector database … the user query and the vectorized context data is sent to the generative AI system which processes the query and returns a generated response which accounts for the information contained in the vectorized context data … the curation system locates and retrieves any available user-defined rules or preferences”).
Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings as in Crabtree with the teachings as in Jin and Loo. The motivation for doing so would ensure the system to have the ability to use system and method disclosed in Crabtree to use multi-modal computing system to combine information from different modalities to create joint embedding spaces; to define user rules; and to obtain plurality of context data may be processed into vectors by an embedding model using user-defined rules thus generating mixed-modality embeddings within a joint embedding space using the mixed-modality data; determining a user-derived constraint for an output application and generating a second coreset of embedding vectors using the coreset of the mixed-modality embeddings and the user-derived constraint in order to determine which modalities are permitted within the mixed-modality summary so that mixed-modality summary can be generated based on the one or more user-derived constraints.
Regarding claim 2, the combination teachings of Jin, Loo, and Crabtree as discussed above also disclose the system of claim 1, wherein: each embedding vector within the second coreset comprises a nearest neighbor joint-modality embedding vector to one of the embedding vectors within the coreset of the mixed-modality embeddings (see Jin, paragraph [0093]: “Embodiments use a transformer network to generate text embeddings and K-Means clustering to identify sentences closest to a centroid of the cluster for summary selection. Some transformer network architectures have objectives that are specific for pre-training. For example, some randomly mask out 10% to 15% of the words in the training data, attempting to predict the masked words, and take in an input sentence and a candidate sentence. Then, the network predicts whether the candidate sentence properly follows the input one. Multiple layers can be used to extract embeddings, where the “cls” layer of the transformer network produces the necessary N×E matrix for clustering, where N is the number of sentences and E is the embeddings dimension. The outputs for other layers in the network produced N×W×E embeddings, where W is equal to the tokenized words”).
The motivation for combining the references has been discussed in claim 1 above.
Regarding claim 6, the combination teachings of Jin, Loo, and Crabtree as discussed above also disclose the system of claim 1, wherein: the user-derived constraint comprises a restriction on a data size for the mixed-modality summary (see Loo, paragraph [0066]: “The inputs can be received from the user computing device 106 and/or the data store 104. The inputs can include, but are not limited to, an original full-size dataset 110, one or more data labels 112A-N corresponding to data in the dataset 110, one or more model types 114A-N for which resulting output from the disclosed techniques can be used for, and/or a data compression size 116 indicating a desired size of the resulting output from the disclosed techniques (e.g., a size of a resulting distilled, synthetic dataset)”).
The motivation for combining the references has been discussed in claim 1 above.
Regarding claim 7, the combination teachings of Jin, Loo, and Crabtree as discussed above also disclose the system of claim 1, further comprising: determining an output device constraint for an output device for outputting the mixed-modality summary, the generating the second coreset of embedding vectors includes generating the second coreset of embedding vectors using the output device constraint (see Loo, paragraph [0122]: “Software instructions, algorithms (e.g., the process(es) for distilling a dataset to a coreset), and data can be coded and stored within the memory for instructing the CPU. Support circuits can also be connected to the CPU for supporting the processor in a conventional manner. The support circuits may include conventional cache, power supplies, clock circuits, input/output circuitry, and/or subsystems, and the like”).
The motivation for combining the references has been discussed in claim 1 above.
Regarding claim 8, the combination teachings of Jin, Loo, and Crabtree as discussed above also disclose the system of claim 7, wherein: the output device constraint comprises a type of output device used for outputting the mixed-modality summary (see Jin, paragraph [0032]: “an IO controller may represent or interact with a modem, a keyboard, a mouse, a touchscreen, or a similar device. In some cases, an IO controller may be implemented as part of processor 205. In some cases, a user may interact with a device via IO controller or via hardware components controlled by an IO controller”); and
the outputting the mixed-modality summary comprises outputting the mixed-modality summary using the output device (see Jin, paragraph [0032]: “I/O module 265 enables a user or networked device to communicate with video summarization apparatus 200. Embodiments of I/O module 265 include an IO controller and/or a user interface. An I0 controller may manage input and output signals for a device”).
The motivation for combining the references has been discussed in claim 1 above.
Regarding claim 9, the combination teachings of Jin, Loo, and Crabtree as discussed above also disclose the system of claim 7, wherein: the mixed-modality data includes text data, image data, audio data, and video data (see Jin, paragraph [0064]: “a method for multimodal unsupervised video temporal segmentation is described. One or more aspects of the method include receiving a video and a transcript of the video; generating visual features representing frames of the video using an image encoder; generating language features representing the transcript using a text encoder, wherein the image encoder and the text encoder are trained based on a correlation between training visual features and training language features; and segmenting the video into a plurality of video segments based on the visual features and the language features. Some examples of the method, apparatus, non-transitory computer readable medium, and system further include converting audio data associated with the video to obtain the transcript”); and
the outputting the mixed-modality summary includes displaying the mixed-modality summary using the output device (see Jin, paragraph [0033]: “A user interface may enable a user (e.g., user 115 with reference to FIG. 1) to interact with a device. In some embodiments, the user interface may include an audio device, such as an external speaker system, an external display device such as a display screen, or an input device (e.g., remote control device interfaced with the user interface directly or through an IO controller module)”).
The motivation for combining the references has been discussed in claim 1 above.
Regarding claim 12, the combination teachings of Jin, Loo, and Crabtree as discussed above also disclose the system of claim 7, wherein: the output device comprises one of a watch, a head-mounted display device, a smartphone, or a laptop computer (see Jin, paragraph [0032]: “I/O module 265 enables a user or networked device to communicate with video summarization apparatus”); and
the mixed-modality summary includes an audio component and a video component (see Jin, paragraph [0071]: “a user provides a video for summarization (e.g., long livestream video). In some cases, the operations of this step refer to, or may be performed by, a user as described with reference to FIG. 1. In some cases, the video includes visual data and audio data, and the audio data is preprocessed to produce a transcript of the video”).
The motivation for combining the references has been discussed in claim 1 above.
Claim 13 is rejected for the same reasons as discussed in claim 1 above. In addition, the combination teachings of Jin, Loo, and Crabtree as discussed above also disclose determining a user-derived constraint for an output application (see Crabtree, paragraph [0387]: “the user-defined rules/preferences may be defined by an entity (e.g., a company). Exemplary rules or preferences can include, but are not limited to, conditional generation preferences, formatting rules, language rules, style rules, geographic rules, environmental rules, and timing rules”); and
generating a second coreset of embedding vectors by remapping at least one embedding vector from the coreset of the mixed-modality embeddings (see Loo, paragraph [0101]: “The KRR on the kernel matrices can be recomputed with the ith individual coreset element removed, with Kx,S\iKS\i,S\i being the resulting kernel matrices with the ith row/column corresponding to the ith coreset entry removed”) , each embedding vector of the second coreset of embedding vectors has fewer modalities than the first number of modalities, the generating the second coreset of embedding vectors includes generating the second coreset of embedding vectors based on the user-derived constraint (see Loo, paragraph [0068]: “the inputs can be previously determined by the user at the user computing device 106 or another relevant user at another computing device. The previously-determined inputs can then be stored at the data store 104 and retrieved by the computer system 102 in block A (120) at another time. In still other instances, in lieu of or in addition to a user determining the inputs, the inputs can be provided by and/or to the user”).
The motivation for combining the references has been discussed in claim 1 above.
Claim 18 is rejected for the same reasons as discussed in claim 9 above.
Claim 19 is rejected for the same reasons as discussed in claim 1 above. In addition, the combination teachings of Jin, Loo, and Crabtree as discussed above also disclose a hardware storage device configured to store mixed-modality data covering data from a first number of modalities (see Jin, paragraph [0073]: “the system can provide the video summary directly to the user through an application, or to a storage such as a database”).
Claims 3-4, 15-16, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Jin, Loo, and Crabtree as applied to claim 1, and further in view of Gori et al. (US 20240144520 A1, hereinafter referred to as “Gori”).
Regarding claim 3, the combination teachings of Jin, Loo, and Crabtree as discussed above disclose all the claimed limitations with the exceptions of the system of claim 1, wherein: the generating the second coreset of embedding vectors includes detecting that a joint-modality embedding that has fewer modalities than the first number of modalities is within a threshold distance of a mixed-modality embedding within the coreset and replacing the mixed-modality embedding with the joint-modality embedding within the second coreset of embedding vectors.
Gori from the same or similar fields of endeavor discloses the system of claim 1, wherein: the generating the second coreset of embedding vectors includes detecting that a joint-modality embedding that has fewer modalities than the first number of modalities is within a threshold distance of a mixed-modality embedding within the coreset and replacing the mixed-modality embedding with the joint-modality embedding within the second coreset of embedding vectors (see Gori, paragraph [0388]: “the scene-based image editing system 106 utilizes the joint embedding space 1816 to manipulate the visual feature maps 1810 with the text instructions of the modification input 1804 a-1804 b via vector arithmetic operations. When manipulating certain objects or object attributes, the object modification neural network 1806 aims to modify only specific regions while keeping other regions unchanged. Accordingly, the object modification neural network 1806 conducts vector arithmetic operations between the visual feature maps 1810 represented as V∈Figure US20240144520A1-20240502-P00012 1024×7×7 and the textual features 1814 a-1814 b (e.g., represented as textual feature vectors”).
Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings as in Gori with the teachings as in Jin, Loo, and Crabtree. The motivation for doing so would ensure the system to have the ability to use system and method disclosed in Gori to utilize the joint embedding space to manipulate the visual feature maps with the text instructions of the modification input via vector arithmetic operations wherein the object modification neural network aims to modify only specific regions when manipulating certain objects or object attributes; to replace the spatial modulation block with an additional global modulation block; to perform a global modulation with regard to the local intermediate vector and to utilize a first global modulation block and a second global modulation block thus generating the modified coreset of embedding vectors by detecting that a joint-modality embedding that has fewer modalities than the first number of modalities is within a threshold distance of a mixed-modality embedding within the coreset; replacing the mixed-modality embedding with the joint-modality embedding within the second coreset of embedding vectors; generating joint-modality embeddings within the joint embedding space wherein each embedding of the joint-modality embeddings has fewer modalities than the first number of modalities; and replacing at least one of the mixed-modality embeddings within the coreset with one of the joint-modality embeddings within the joint embedding space in order to embedding two or more modalities within a joint embedding space so that mixed-modality embeddings can be generated using the mixed-modality data within a joint embedding space.
Regarding claim 4, the combination teachings of Jin, Loo, Crabtree, and Gori as discussed above also disclose the system of claim 1, further comprising: generating joint-modality embeddings within the joint embedding space, each embedding of the joint-modality embeddings has fewer modalities than the first number of modalities (see Gori, paragraph [0388]: “the scene-based image editing system 106 utilizes the joint embedding space 1816 to manipulate the visual feature maps 1810 with the text instructions of the modification input 1804 a-1804 b via vector arithmetic operations. When manipulating certain objects or object attributes, the object modification neural network 1806 aims to modify only specific regions”); and
replacing at least one of the mixed-modality embeddings within the coreset with one of the joint-modality embeddings within the joint embedding space (see Gori, paragraph [0214]: “the scene-based image editing system 106 utilizes a global modulation block followed by another global modulation block. For example, the scene-based image editing system 106 replaces the spatial modulation block 603 with an additional global modulation block. In such an embodiment, the scene-based image editing system 106 replaces APN (and spatial tensor) and corresponding spatial modulation illustrated in FIG. 6 with a skip connection. For example, the scene-based image editing system 106 utilizes the global intermediate feature to perform a global modulation with regard to the local intermediate vector. Thus, in some cases, the scene-based image editing system 106 utilizes a first global modulation block and a second global modulation block”).
The motivation for combining the references has been discussed in claim 3 above.
Claim 15 is rejected for the same reasons as discussed in claim 3 above.
Claim 16 is rejected for the same reasons as discussed in claim 4 above.
Claim 20 is rejected for the same reasons as discussed in claim 3 above. In addition, the combination teachings of Jin, Loo, Crabtree, and Gori as discussed above also disclose the user-derived constraint comprises a restriction on a data size for the mixed-modality summary (see Loo, paragraph [0066]: “The inputs can be received from the user computing device 106 and/or the data store 104. The inputs can include, but are not limited to, an original full-size dataset 110, one or more data labels 112A-N corresponding to data in the dataset 110, one or more model types 114A-N for which resulting output from the disclosed techniques can be used for, and/or a data compression size 116 indicating a desired size of the resulting output from the disclosed techniques (e.g., a size of a resulting distilled, synthetic dataset)”).
The motivation for combining the references has been discussed in claim 1 above.
Claims 10-11 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Jin, Loo, and Crabtree as applied to claim 7, and further in view of Bikumala et al. (US 20200387640 A1, hereinafter referred to as “Bikumala”).
Regarding claim 10, the combination teachings of Jin, Loo, and Crabtree as discussed above disclose all the claimed limitations with the exceptions of the system of claim 7, further comprising: detecting that an amount of noise within an operating environment of the output device is greater than a threshold level of noise and preventing an audio component from being a part of the mixed-modality summary in response to detecting that the amount of noise within the operating environment of the output device is greater than the threshold level of noise.
Bikumala from the same or similar fields of endeavor discloses the system of claim 7, further comprising: detecting that an amount of noise within an operating environment of the output device is greater than a threshold level of noise and preventing an audio component from being a part of the mixed-modality summary in response to detecting that the amount of noise within the operating environment of the output device is greater than the threshold level of noise (see Bikumala, paragraph [0027]: “The computing device 102 may include multiple sensors, such as, for example, a microphone 112, a camera 114, and the like. The sensors 112, 114 may provide sensor data to the computing device 102 to enable the computing device 102 to determine whether the computing device 102 is located in a public environment or a private environment. For example, if the audio data from the microphone 112 indicates that there is a large amount of background noise, e.g., the audio data satisfies a noise threshold (e.g., at least 35, 40, 45, 50 decibels (db) or the like), then the computing device 102 may determine that the environment is a public environment”).
Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings as in Bikumala with the teachings as in Jin, Loo, and Crabtree. The motivation for doing so would ensure the system to have the ability to use system and method disclosed in Bikumala to provide sensor data to the computing device to enable the computing device to determine whether the computing device is located in a public environment or a private environment by detecting if the audio data satisfies a noise threshold; and to determine what to display by comparing the display data size with a screen size of the display device and to display data when the display device has a size greater than a threshold size thus detecting that an amount of noise within an operating environment of the output device is greater than a threshold level of noise and preventing an audio component from being a part of the mixed-modality summary in response to detecting that the amount of noise within the operating environment of the output device is greater than the threshold level of noise and detecting that a display size for the output device is less than a threshold display size and preventing a video component from being a part of the mixed-modality summary in response to detecting that the display size for the output device is less than the threshold display size in order to determine output device constraints so that a coreset associated with a mixed-modality summary can be generated based on output device constraints such as the threshold of level of noise and display size.
Regarding claim 11, the combination teachings of Jin, Loo, Crabtree, and Bikumala as discussed above also disclose the system of claim 7, further comprising: detecting that a display size for the output device is less than a threshold display size and preventing a video component from being a part of the mixed-modality summary in response to detecting that the display size for the output device is less than the threshold display size (see Bikumala, paragraph [0041]: “Depending on a screen size of the display device 104, more than one entry wheel may be displayed. For example, as illustrated in FIG. 4, an alphabetic entry wheel 402 may be displayed, a special character entry wheel 404 may be displayed, a numeric entry wheel 406 may be displayed, or any combination thereof. The entry wheels 402, 404, 406 may be displayed when the display device 104 has a size greater than a threshold size”).
The motivation for combining the references has been discussed in claim 10 above.
Claim 17 is rejected for the same reasons as discussed in claim 10 above.
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Jin, Loo, and Crabtree as applied to claim 13, and further in view of Peng et al. (US 20190325084 A1, hereinafter referred to as “Peng”).
Regarding claim 14, the combination teachings of Jin, Loo, and Crabtree as discussed above also disclose the method of claim 13, further comprising: each embedding vector within the second coreset comprises a nearest neighbor joint-modality embedding vector to one of the embeddings within the coreset of the mixed-modality embeddings (see Jin, paragraph [0093]: “Embodiments use a transformer network to generate text embeddings and K-Means clustering to identify sentences closest to a centroid of the cluster for summary selection. Some transformer network architectures have objectives that are specific for pre-training. For example, some randomly mask out 10% to 15% of the words in the training data, attempting to predict the masked words, and take in an input sentence and a candidate sentence. Then, the network predicts whether the candidate sentence properly follows the input one. Multiple layers can be used to extract embeddings, where the “cls” layer of the transformer network produces the necessary N×E matrix for clustering, where N is the number of sentences and E is the embeddings dimension. The outputs for other layers in the network produced N×W×E embeddings, where W is equal to the tokenized words”).
Regarding claim 14, the combination teachings of Jin, Loo, and Crabtree as discussed above disclose all the claimed limitations with the exceptions of the user-derived constraint comprises a restriction on a length of time for the mixed-modality summary.
Peng from the same or similar fields of endeavor discloses the user-derived constraint comprises a restriction on a length of time for the mixed-modality summary (see Peng, paragraph [0070]: “The one or more parameters may determine one or more of a length of the summary of each content object, a topic of the summary of each content object, a modality of the summary of each content object, other suitable properties of each content object, or any combination thereof. In particular embodiments, the input from the first user may comprise both explicit and implicit signals”).
Therefore it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize the teachings as in Peng with the teachings as in Jin, Loo, and Crabtree. The motivation for doing so would ensure the system to have the ability to use system and method disclosed in Peng to receive a user request for a summarization of a particular type of content objects from a client system associated with a first user; to determine one or more modalities associated with the user request wherein the one or more parameters may determine one or more of a length of the summary of each content object thus comprising a user-derived constraint on a length of time for the mixed-modality summary in order to generate the mixed-modality summary based on the user request so that the length of mixed-modality summary cannot exceed user defined total length of time.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NIENRU YANG whose telephone number is (571)272-4212. The examiner can normally be reached Monday-Friday 10AM-6PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, THAI TRAN can be reached at 571-272-7382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
NIENRU YANG
Examiner
Art Unit 2484
/NIENRU YANG/Examiner, Art Unit 2484
/THAI Q TRAN/Supervisory Patent Examiner, Art Unit 2484