DETAILED ACTION
1. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
2. The following office action is a Final Office Action in response to the communications received on 05/21/2026.
Claims 1, 8, 15, 21 and 22 have been amended; claims 4-7, 11-14, 17-20 and 23 have been canceled; and new claims 24-26 have been added. Therefore, currently claims 1-3, 8-10,15, 16, 21, 22 and 24-26 are pending in this application.
Claim Rejections - 35 USC § 101
3. Non-Statutory (Directed to a Judicial Exception without an Inventive Concept/Significantly More)
35 U.S.C.101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
● Claims 1-3, 8-10,15, 16, 21, 22 and 24-26 are rejected under 35 U.S.C.101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1
The current claims fall within one of the four statutory categories of invention (MPEP 2106.03).
Step 2A [Wingdings font/0xE0] Prong One:
The claim(s) recite a judicial exception, namely an abstract idea, as shown below:
— Considering each of claims 1, 8 and 15 as representative claims, the following claimed limitations recite an abstract idea (note that the term “video” is construed as a “content item” since a video is itself a content item):
receive a request for a [content item] illustrating an event at a target location and a target time frame;
determine one or more requirements for the [content item];
[collect], based on the one or more requirements, a plurality of crowdsourced [content items] illustrating the event;
extract a set of metadata from the plurality of crowdsourced [content items], the set of metadata including geographical information, capture orientation, and timestamp information;
cluster the plurality of crowdsourced [content items] into location clusters based on a geographical proximity defined by a pre-set threshold distance from the target location;
determine, based on the metadata, a chronological sequence of the clustered [content items] within the location cluster during the target time frame;
determine a first temporal image set of the plurality of crowdsourced [content items];
synchronize the first temporal image set in terms of temporal alignment and spatial alignment based on the metadata;
construct a three-dimensional (3D) model of the event from the synchronized first temporal image set;
create [content item] of the event, the [content item] including sequentially arranging multiple 3D models from a plurality of temporal image sets including the first temporal image set, the [content item] meeting the one or more requirements.
Thus, the limitations identified above recite an abstract idea since the limitations correspond to certain methods of organizing human activity and/or mental processes, which are part of the enumerated groupings of abstract ideas identified according to the current eligibility standard (see MPEP 2106.04(a)).
The claims correspond to certain methods of organizing human activity, wherein a content item—such as a video—is created, in response to a request received from a user, including on one or more requirements; and wherein the content item is created based on analyzing information gathered from content items that other users provided, including: identifying geographical information; capture orientation, timestamp information; grouping the content items based on location—such as, a geometrical proximity defined by a preset threshold distance from the target location; determining, using the metadata, the chronological sequence of the content items in the group above; determining a first image set of the plurality of crowdsourced content using a mathematical formula; synchronizing the first temporal image set in terms of temporal alignment and spatial alignment based on the metadata; constructing a 3D model of the event from the synchronized first temporal image set; and creating the content item of the event, the content item including sequentially arranging multiple 3D models from a plurality of temporal image sets, including the first image set above, etc.
Note that the finding above—certain methods of organizing human activity—is consistent with the specification. For instance, the specification describes that a user provides a set of content, including a request to create media-based case study for an educational topic; and wherein, once a requirement(s) is determined for the educational topic, a search query is constructed, so that one or more relevant media content items are identified—using the search query—from media items that other users provided; and thereby, a customized content item (a customized video) is created by considering one or more factors, including: the location of capture, the direction of capture, the timing of the capture, etc., and eventually the end result is presented to the user (e.g., see [0031] to [0035]).
The current claims also correspond to a mental process; such as, a process that can be performed in the human mind and/or using a pen and paper, given the limitations that recites the process of: determining one or more requirements for a given content item (e.g. a video); extracting metadata that includes multiple pieces of information (e.g., geographical information, capture orientation, timestamp information) ; and determining, based on temporal information, a chronological sequence of clustered content items (e.g., videos) within the location clusters during the target time frame, etc.
Thus, given such limitations that recite the process of evaluation, observation, and/or judgement, etc., the current claims do recite a mental process.
Similarly, given the limitation that recites a mathematical formula; namely, the claimed “linear interpolation function” (also see [0051] of the specification, which signifies the “linear interpolation formula”), each of the current claims further recites a mathematical concept.
Step 2A [Wingdings font/0xE0] Prong Two:
The claim(s) recite additional element(s), wherein a computer-based system is utilized as a tool to facilitate the recited steps/functions regarding: receiving input or a request (e.g., “receiving, by a content creation program, a request for a volumetric video illustrating an event at a target location and target time frame”); determining one or more requirement(s) (e.g., “determining, by the content creation program, one or more requirements for the volumetric video”); collecting content items from sources over a network (e.g., “retrieving, based on the one or more requirements, a plurality of crowdsourced videos illustrating the event from a social media network”); analyzing the collected content items (e.g., “extracting, by the content creation program, metadata from the plurality of crowdsourced videos . . . clustering the plurality of crowdsourced videos into location clusters based on a geographical proximity defined by a pre-set threshold distance from the target location”); determining one or more results (e.g., “determining, based on the metadata, a chronological sequence of clustered videos within the location clusters during the target time frame; determining a first temporal image set of the plurality of crowdsourced videos by applying geometric operations comprising: joining multiple points corresponding to the plurality of crowdsourced videos . . . using a linear interpolation function from a first point to subsequent points and connecting a last point back to the first point to define a closed loop contour”); synchronizing images (e.g., “synchronizing the first temporal image set in terms of temporal alignment and spatial alignment based on the metadata”); generating a customized content (e.g., “constructing a three-dimensional (3D) model of the event from the synchronized first temporal image set”; “creating the volumetric video of the event, the volumetric video including sequentially arranging multiple 3D models from a plurality of temporal image sets including the first temporal image set, the volumetric video meeting the one or more requirements”), etc.
However, the claimed additional element(s) fail to integrate the abstract idea into a practical application since the additional element(s) are utilized merely as a tool to facilitate the abstract idea. Thus, when each claim is considered as a whole, the additional element(s) fail to integrate the abstract idea into a practical application since they fail to impose meaningful limits on practicing the abstract idea. For instance, when each of the claims is considered as a whole, none of the claims provides an improvement over the relevant existing technology.
The observations above confirm that the claims are directed to an abstract idea.
Step 2B
Accordingly, when the claim(s) is considered as a whole (i.e., considering all claim elements both individually and in combination), the claimed additional elements do not provide meaningful limitations to transform the abstract idea into a patent eligible application of the abstract idea such that the claim(s) amounts to “significantly more” than the abstract idea itself (also see MPEP 2106). The claimed additional elements are directed to conventional computer elements, which are serving merely to perform conventional computer functions. Thus, none of the claims, considered as a whole, recites an element—or a combination of elements—directed to an inventive concept.
It is also worth to note that use of the conventional computer/network technology to facilitate the process of generating pertinent content to a user (e.g., one or more customized videos, etc.), based on a request and/or parameters received from the user, etc., is already directed to a well-understood, routine or conventional activity in the art (see US 2016/0105708; US 2015/0143239; US 2013/0117780, etc.).
Of course, the same is true regarding the process of identifying, using metadata, related scenes from one or more videos; including generating a customized video based on compiling the identified scenes, etc. (e.g., see US 2016/0080835; US 2013/0325869; US 2009/0132924, etc.).
The above observation confirms that each of the current claims fails to amount to “significantly more” than an abstract idea.
It is worth noting that the above analysis already encompasses each of the current dependent claims (i.e., claims 2, 3, 9, 10, 13, 16, 21, 22 and 24-26). Particularly, each of the dependent claims also fails to amount to “significantly more” than the abstract idea since each dependent claim is directed to a further abstract idea, and/or a further conventional computer element utilized to facilitate the abstract idea.
Thus, the findings above demonstrate that none of the claims is implementing an element—or a combination of elements—directed to an inventive concept (e.g., none of the current claims is reciting an element—or a combination of elements—that provides a technological improvement over the existing/conventional technology).
► Applicant’s arguments directed to section §101 have been fully considered (the response filed on 05/21/2026). However, the arguments art not persuasive at least for the following reasons:
Firstly, while referring to the current amendment, Applicant asserts, “claim 1 is directed to patent-eligible subject matter under 35 U.S.C. § 101 because it recites a specific technological improvement in computer-based video processing and reconstruction, consistent with Federal Circuit precedent, including McRO, Inc. v. Bandai Namco Games America Inc . . . applying defined rules to automate a process, previously performed manually, constitutes a technological improvement and is not an abstract idea. Here, amended claim 1 similarly recites specific rule-based operations, including: representing each crowdsourced video as a point defined by location and a direction vector; interpolating points using a linear interpolation function; and connecting a final point back to an initial point to form a closed-loop contour. These limitations define how the volumetric video is generated, not merely the result, and thus parallel the rule-based animation techniques found eligible in McRO” (emphasis added).
However, Applicant appears to misapply the court’s finding regarding McRO. It is again worth noting that McRO is providing a technological improvement in computer animation (e.g., automatic lip synchronization and facial expression animation), wherein the system implements specific rules that improved existing technology. In contrast, the process automatically constructing—from one or more videos—a customized video based on one or more attributes, is already part of the existing computer/network technology. Although a reference is not necessarily required to demonstrate the fact above, one or more of the references cited as part of the Step 2B analysis already confirm the fact above. For instance, Von Sneidern (US 2016/0080835) describes such a system directed to the existing computer/network technology; and this system automatically generates—from one or more videos—a customized video based on one or more desired attributes (e.g., see [0023] to [0026], emphasis added),
“. . . users often may record any and everything with the hopes of editing out the less interesting portions at a later time. This practice may lead to hours of footage to edit, which can be a daunting task for many users. Often, these hours of video are never actually edited leaving viewers with hours of video to sort through to find the most interesting moments”
“. . . the present disclosure address these and other shortcomings by providing methods and/or systems for automatically creating a compilation video from one or more source videos. The compilation video . . . may highlight many of the interesting parts of the one or more source videos while filtering out the less interesting parts. Techniques described herein may be used to identify and learn what makes a video ‘interesting’ and then that knowledge may be used to generate the compilation video”
“ A compilation video is a video that includes more than one video clip selected from portions of one or more source video(s) and joined together to form a single video. A compilation video may be created based on the metadata associated with the source videos. Compilation videos may further be created based on relevance scores assigned to video frames and/or video clips. A relevance score may indicate, for example, a level of interestingness of the content in a video clip, which may include a level of excitement occurring with the source video as represented by motion data, the location where the source video was recorded, the time or date the source video was recorded . . .”
“. . . Metadata of a video may include one or more features. These features may include any data that is captured in association with the recording of the video, such as geo location, motion of the video capturing device, etc. . . .”
Note also that the teaching of Von Sneidern above applies to different types of videos, including 3D videos, since the reconstruction of 3D videos based on one or more attributes is already part of the existing computer/network technology. For instance, Birkbeck (US 2015/0143239) describes an earlier system that reconstructs, from one or more videos sources, a 3D video bases on the analysis of various attributes (see [0029], [0052], [0058], [0059], [0063], emphasis added),
“. . . video discovery module 202 scans media items 242 and identifies media items having metadata or other cues that identify the real-world event . . . The metadata cues may include information in the title or description of the video, user provided or system generated tags or categories, date and time information associated with the media items, geolocation information (e.g., GPS data) associated with the media items, or other information. Upon determining that a particular media item 242 is associated with a given real-world event, video discovery module may add the media item 242 to an event list 244 corresponding to the real-world event”
“ The theory of multiple view geometry provides the mathematical tools to do reconstruction of camera poses and scene geometry from image-derived point correspondences. Although work has been done on 3D constructions from multiple camera views, many techniques only work with assumptions that the internal calibrations (e.g., focal lengths, principal points) for the cameras are known . . . the system uses the pure camera rotation present in the user generated videos to automatically extract the internal calibration”
“ For computational reasons, the system may first reduce the number of frames input to the reconstruction, by selecting only a few salient frames from each video sequence by considering the number of features, quality of each frame, and amount of temporal motion. Once the system has selected images for reconstruction, it can extract SIFT features from each image and match pairs of images using these features . . .”
“. . . Once we have the two-view models, the system can iteratively add two-view models together and do bundle adjustment, to get the final 3D model containing all the cameras”
“ After solving for the optimization, a combined edited video may be created . . .”
Accordingly, unlike Applicant’s assertion, none of the current claims is analogous to McRO, which automated a process that previously performed manually by human artists. This is again because it is already part of the existing computer/network technology to automatically generate a customized video from one or more videos gathered from one or more sources. Of course, depending on the specific type of customization, one may specify one or more desired attributes (e.g., geographical location, capture orientation, timestamp orientation, chronological order, etc.), which the algorithm is required to use when constructing the customized video. Similarly, one may also use one or more existing algorithms to accomplish the customization process above. However, such existing computer implementation that uses one or more attributes for the customization process, including one or more existing algorithms to accomplish the customization, etc., does not constitute a technological improvement.
In particular, when considering Applicant’s claimed (and disclosed) implementation, neither the current claims nor the original specification is demonstrating a feature—or a combination of features—that provides a technological improvement. Instead, the current claims, including the original specification, are describing the procedures that the claimed (or the disclosed) system/method chose to use to build the customized video. For instance, per the original disclosure, a mathematical concept is utilized to create a closed loop contour by joining multiple points, wherein each point is defined mathematically to include a location and a direction; and subsequently, each point is interpolated sequentially using one of the existing algorithms—namely, a linear interpolation formula, in order to connect two or more points, including connecting the last point to the first point ([0051]). The system further synchronizes the content in terms of time and spatial alignment ([0052]); and subsequently, the system (i) reconstructs the 3D model using one or more existing techniques—e.g., Structure from Motion, Multi-View Stereo ([0053]), (ii) extracts, using an attribute (e.g., silhouette of a subject), closed loop contour information from the 3D model above ([0054]), (iii) verifies the consistency of each closed loop contour across the source videos using one or more existing techniques (e.g., stereo matching technique, a depth sensor, or a depth estimation method) ([0055]); and finally, (iv) converts the above 3D model into a volumetric video ([0056]).
Accordingly, except for describing the steps being performed to generate the customized video, including one or more of the existing algorithms that the system is using to accomplish the customization process, the specification does not demonstrate a feature—or a combination of features—that is directed to a technological improvement. Thus, Applicant’s alleged technological improvement is not persuasive.
Secondly, Applicant is asserting that “claim 1 define a technical pipeline for transforming raw media data into a structured 3D representation, which is an improvement to a specific technological process. That is, the claims represent an improvement over conventional volumetric video systems, which typically rely on controlled multi-camera setups, by enabling reconstruction from heterogeneous, uncalibrated video inputs . . . amended claim 1 recites a specific set of rule-based computational steps that improve the technical process of generating volumetric video from crowdsourced media, analogous to the rule-based animation techniques found patent-eligible in McRO. Claims 8 and 15 recite similar claim language . . . claims 1, 8, and 15 are streamlined here by removing non-essential data analysis language (e.g., object identification) and further emphasizing the rule-based geometric operations that generate the 3D structure. This aligns the claim squarely with McRO by focusing on how the result is achieved through defined computational rules” (emphasis added).
However, besides misconstruing the transformation test, Applicant also appears to mischaracterize the conventional computer/network technology. For instance, the transformation test evaluates whether an article (e.g., a physical substance) is being transformed from one state into a different sate or thing—such as, transforming a raw, uncured rubber, into to a precision-molded rubber, see Diehr, 450 U.S. at 184, 209 USPQ at 21.
In contrast, Applicant’s claimed—and disclosed—implementation is merely constructing or customizing multimedia information (e.g., a customized 3D video) from two or more pieces of multimedia information (e.g., two or more videos gathered from one or more online sources). Thus, there is no transformation from one state to a different sate or thing since both the initial and final states are the same (i.e., both states are multimedia information in the form of video). Consequently, Applicant’s arguments are not persuasive.
In addition, again unlike Applicant’s assumption, the conventional technology is not limited merely to controlled multi-camera setups. In fact, the evidence presented above (e.g., see the cited sections above) effectively invalidates Applicant’s assumptions. This is because the conventional technology considers various attributes (e.g., timing information related to events, location information related to events, one or more audio and/or image cues detected within in the videos, etc.) to identify and categorize relevant video pieces in order to construct the customized video. Thus, Applicant’s conclusory assertion directed to the conventional technology is also not persuasive.
Note also that Applicant’s repetitive argument, which attempts to correlate the current claims with that of McRO, is once again not persuasive since none of the current claims is analogous to McRO. In particular, unlike the case of McRO, the discussion presented above already demonstrates that Applicant’s claimed and/disclosed implementation is directed to the existing computer/network technology; namely, the process of automatically generating a customized 3D video based on the analysis of one or more attributes (again see the discussion presented above). In this regard, even the specification confirms the one or more existing techniques, which the disclosed system is utilizing to reconstruct the 3D volumetric video (see above the citations presented from the specification). In contrast, McRO implemented specific rules that improved existing technology; namely, an improvement in computer animation. Consequently, Applicant’s arguments directed to McRO are once again not persuasive.
Thus, at least for the reasons discussed above, the Office concludes that none of the current claims, when considered as a whole, implements an inventive concept that amounts to “significantly more” than an abstract idea.
Prior Art
● Similar to the point made in the previous office action, when each of the current claims is considered as a whole, the prior art does not teach or suggest the current claims (regarding the state of the prior art, see the office-action dated 05/21/2025).
Conclusion
Applicant’s amendment necessitated the new grounds of rejection presented in this final office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filled within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRUK A GEBREMICHAEL whose telephone number is (571) 270-3079. The examiner can normally be reached from 7:00 AM - 3:00 PM.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, PETER VASAT can be reached on (571) 270-7625. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/BRUK A GEBREMICHAEL/Primary Examiner, Art Unit 3715