DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Compact Prosecution
With respect to Claim Interpretation, the Examiner has provided some notes regarding “[BRI on the record]” throughout the Office Action, so that the record is clear about the scope of the claimed invention, and the record is also clear about the basis for the Examiner’s analyses. A clear record of the claim interpretation could expedite the examination by creating the condition to allow the examination to focus on Applicant’s inventive concept and its comparison with related prior art.
If there are disagreements, Applicant may present an alternative interpretation based on MPEP 2111. The Examiner will adopt Applicant’s interpretation on the record, if Applicant’s interpretation is reasonable and/or arguments are persuasive.
Applicant may amend claims relying on the Examiner’s claim interpretation provided on the record.
Priority
The Office failed to electronically retrieve, under the priority document exchange program, the foreign application KR10-2023-0142088 to which priority is claimed on 3/23/2025.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-3, 7, 9-11, and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (“Deep 6-DoF camera relocalization in variable and dynamic scenes by multitask learning”) (thereafter as Wang-1), and further in view of Seo et al. (US 20230401774 A1).
Regarding Claim 1, Wang-1 teaches An asset creation method (
[BRI on the record] With respect to “asset,” the Examiner is reading the limitation to mean: something of value based on the plain meaning of the term.
[Mapping Analyses]
“By the above procedure, we can generate the localization dataset with changeable objects. In the paper, three scenes are gathered and constructed, composed of an office, a sitting room and a bedroom.” Wang-1 p. 4 right col.
PNG
media_image1.png
802
762
media_image1.png
Greyscale
), comprising:
simultaneously performing segmentation (segmentation to generate mask image in Wang-1 Fig. 1) and position information identification (e.g., point cloud in Wang-1 Fig. 1) on a target object (the target scene, e.g., a room) to be assetized (constructed scene (camera pose + mask + point cloud), which are useful for generating scenes depicted in Wang-1 Fig. 8) from a video received from a user terminal (“As illustrated in Fig. 2, the whole process takes the images by a mobile phone as inputs and . . ..” Wang-1 p. 4 right col.) based on a parallel network (Wang-1 Fig. 11 showing a parallel network with parallel components) including a three-dimensional (3D) semantic segmentation network (“… MMLNet+ outputs a semantic segmentation branch.” Wang-1 p. 2 right column.) and a Long Short-Term Memory network (LSTM layers in Wang-1 Fig. 3) (
[BRI on the record] With respect to “simultaneously,” the Examiner is reading the limitation to mean: the claimed method performs both the segmentation and position information identification. The interpretation is in light of Applicant’s disclosure. Applicant’s fig. 6:
PNG
media_image2.png
608
434
media_image2.png
Greyscale
, where S640 “position identification” occurs after S610 “semantic segmentation for target area.”
With respect to 3D “semantic segmentation,” the Examiner is reading the limitation to mean: semantic segmentation for 3D features.
[0060] When the user terminal records a video and transmits the video to the asset creation apparatus, the video data preprocessor may convert the video into 3D video data in a point cloud format or the like.
Spec. ¶ 60.
With respect to “a parallel network including A, B, and C,” under BRI, the limitation does not require A, B, and C are in parallel with each other.
[Mapping Analyses]
Wang-2 teaches three-dimensional (3D) segmentation, stating “In detail, the multitask output covers camera pose, 3D point cloud and segmentation, allowing MMLNet and MMLNet+ to learn relations between 2D images and 3D poses more sufficiently. To share the features among three tasks, the feature fusion block is designed.” Wang-2 p. 2 right col. “It should be noticeable that our process leverages the projection of marked 3D points to construct mask images instead of labeling directly.” Wang-2 p. 4 right col.
Wang-2 teaches its model could use videos, e.g., testing video, as input, stating “As illustrated in Fig. 2, the whole process takes the images by a mobile phone as inputs and . . .” (Wang-1 p. 4 right col.) and “Meanwhile, due to the testing video, the accumulative error problem of SLAM systems is not obvious [as an advantage over Wang-2’s model]” (Wang-1 p. 12 left col.).
The Examiner’s secondary reference also teaches the use of video as input.); and
generating a 3D video feature (changeable objects contain the books, tea caddy, tissue, mouse, keyboard and notebook, correctly localized in a constructed 3D environment) of the target object (target scene) from results of the segmentation and the position information identification constructed 3D scene of = camera pose + mask + 3D point cloud; or virtual object, correctly placed on a constructed desk and in correct spatial relationships with other constructed objects on the desk in Fig. 8) in conformity with the 3D video feature (the constructed 3D scene that contains the changeable object(s)) (
“Furthermore,we define the feature fusion block to implement the feature sharing among three tasks, further promoting the performance in dynamic and variable environments. Finally, experiments on static, dynamic and our constructed variable datasets demonstrate stateof-the-art relocalization performances of MMLNet and MMLNet+.” Wang-1 Abstract. A temporal sequence of images, a video, captures a dynamic environment with changeable objects.
“By the above procedure, we can generate the localization dataset with changeable objects. In the paper, three scenes are gathered and constructed, composed of an office, a sitting room and a bedroom. In each scene, we select some common objects without fixed positions as changeable objects. For example, in the office scene, changeable objects contain the books, tea caddy, tissue, mouse, keyboard and notebook.” Wang-1 p. 4, right col.
“After obtaining camera poses, AR applications always register virtual objects in the real scene. The rendering outcomes are demonstrated in Fig. 8 through evaluated poses by MMLNet+.” Wang-1 p. 11, right col.
PNG
media_image3.png
508
752
media_image3.png
Greyscale
).
Wang-1 does not explicitly disclose; however, Seo teaches generating a 3D video feature of the target object from results of the segmentation and the position information identification based on a covariance matrix (
“In addition, the present invention provides the full-body integrated motion capture method, wherein, in the step (b2), the projection image for the three-dimensional object may be acquired by projecting the point cloud of the three-dimensional object onto a horizontal plane, obtaining a front direction vector on a two-dimensional plane onto which the point cloud is projected, and rotating the three-dimensional object such that a front surface of the three-dimensional object is arranged on a vertical axis by using the front direction vector.” Seo ¶ 23.
“In addition, the present invention provides the full-body integrated motion capture method, wherein, in the step (b2), the front direction vector may be obtained through principal component analysis, in which a covariance matrix may be obtained for the two-dimensional plane onto which the point cloud is projected, an eigenvector for the obtained matrix may be obtained, and a vector having a smallest eigenvalue may be set as the front direction vector.” Seo ¶ 24.).
Seo also teaches the use of video as input (“The full-body integrated motion capture method includes: (a) receiving a multiview color-depth video and a high-resolution video; . . ..” Seo Abstract.).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Seo’s covariance matrix with Wang-1. One of ordinary skill in the art would be motivated to match 3d information with 2D information to capture/reconstruct 3D objects. “In addition, the present invention provides the full-body integrated motion capture method, wherein, in the step (b2), the projection image for the three-dimensional object may be acquired by projecting the point cloud of the three-dimensional object onto a horizontal plane, obtaining a front direction vector on a two-dimensional plane onto which the point cloud is projected, and rotating the three-dimensional object such that a front surface of the three-dimensional object is arranged on a vertical axis by using the front direction vector.” Seo ¶ 23. The goal is consistent with Wang-1’s teaching. “It should be noticeable that our process leverages the projection of marked 3D points to construct mask images instead of labeling directly.” Wang-2 p. 4 right col.
Regarding Claim 2, Wang-1 in view of Seo teaches The asset creation method of claim 1, wherein creating the asset comprises:
calculating a vector pointing from one point to an additional point based on sequence data extracted from 3D video data of the video (
Wang-1:
PNG
media_image4.png
516
768
media_image4.png
Greyscale
Here the predicted trajectory captures the vector change from one camera position (one point) to next camera position (additional point) based on the sequence data of camera positions.
“To solve the problem, we set several fixed spots in the scene during gathering different sequences. By camera poses of settled spots, relative camera transformations are calculated, making all camera poses of images under the same coordinate space.” Wang-1 p. 4 right col. ).
Regarding Claim 3, Wang-1 in view of Seo teaches The asset creation method of claim 2, wherein calculating the vector comprises:
measuring similarity (comprising front direction vector, indicating the level of similarity, e.g., when the vector=0, similarity in the front direction is 100%) to the additional point (next point) while rotating the 3D video data (3D object) from the one point (current point) at a preset angle (present according to covariance matrix) using the covariance matrix (
“In addition, the present invention provides the full-body integrated motion capture method, wherein, in the step (b2), the projection image for the three-dimensional object may be acquired by projecting the point cloud of the three-dimensional object onto a horizontal plane, obtaining a front direction vector on a two-dimensional plane onto which the point cloud is projected, and rotating the three-dimensional object such that a front surface of the three-dimensional object is arranged on a vertical axis by using the front direction vector.” Seo ¶ 23.
“In addition, the present invention provides the full-body integrated motion capture method, wherein, in the step (b2), the front direction vector may be obtained through principal component analysis, in which a covariance matrix may be obtained for the two-dimensional plane onto which the point cloud is projected, an eigenvector for the obtained matrix may be obtained, and a vector having a smallest eigenvalue may be set as the front direction vector.” Seo ¶ 24.).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Seo’s covariance matrix with Wang-1. One of ordinary skill in the art would be motivated to match 3d information with 2D information to capture/reconstruct 3D objects. See. Seo ¶ 23.
Regarding Claim 7, Wang-1 in view of Seo teaches The asset creation method of claim 3, wherein calculating the vector further comprises: performing dimension reduction on a point cloud corresponding to the 3D video data in consideration of a computing resource (
[BRI on the record] With respect to “in consideration of a computing resource,” the limitation does not change the operative step of the method or the structure of an apparatus. Here, “in consideration of” is differentiated from “based on,” and “in consideration of” could be thinking about a factor without taking and action based on the thinking.
[Mapping Analysis]
“In addition, the present invention provides the full-body integrated motion capture method, wherein, in the step (b2), the projection image for the three-dimensional object may be acquired by projecting the point cloud of the three-dimensional object onto a horizontal plane, obtaining a front direction vector on a two-dimensional plane onto which the point cloud is projected, and rotating the three-dimensional object such that a front surface of the three-dimensional object is arranged on a vertical axis by using the front direction vector.” Seo ¶ 23.
When a 3D point cloud is projected onto a plane, dimension reduction is performed, and the point cloud with spatial-temporal changes is mapped to the 3D video data. This mapping is consistent with the specification, which states “. . . the video data preprocessor may convert the video into 3D video data in a point cloud format or the like.” Spec. ¶ 60.
“In addition, the present invention provides the full-body integrated motion capture method, wherein, in the step (b2), the front direction vector may be obtained through principal component analysis, in which a covariance matrix may be obtained for the two-dimensional plane onto which the point cloud is projected, an eigenvector for the obtained matrix may be obtained, and a vector having a smallest eigenvalue may be set as the front direction vector.” Seo ¶ 24.
In addition, Seo teaches saving computing resource, stating “The present invention may capture all movements of the face, the body, and the hand through one operation, so that a work process may be remarkably shortened, and manpower/time/production costs in a VFX field where the motion capture is utilized may be reduced. In addition, when videos are acquired from various angles by using a plurality of cameras, there is an advantage in that a movement may be captured more precisely and in detail.” Time is a measurement for the extent of computing resources used.).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Seo’s covariance matrix with Wang-1. One of ordinary skill in the art would be motivated to match 3d information with 2D information to capture/reconstruct 3D objects. Seo ¶ 23. The goal is consistent with Wang-1’s teaching. “It should be noticeable that our process leverages the projection of marked 3D points to construct mask images instead of labeling directly.” Wang-2 p. 4 right col.
Claims 9-11 and 15 are substantially similar to Claims 1-3 and 7. The rejections analyses based on Wang-1 in view of Seo for Claims 1-3 and 7 are applied to Claims 9-11 and 15. In addition, Claim 9 recites, “An asset creation apparatus, comprising: a processor configured to . . .; and a memory configured to store the video” (Seo Fig. 1; Seo ¶¶ 48-49, 53. Seo discloses storing videos on “a storage medium of the computer terminal.” Seo ¶ 53. Seo does not explicitly disclose that the storage medium could be memory. The Examiner takes an Official Notice that it would have been well-known in the art the memory could be used as a storage medium of a computer to store data. The benefits of combining this well-known knowledge would have been fast access or retrieval of data.).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Seo’s teaching on using computers with Wang-1. One of ordinary skill in the art would be motivated to achieve faster and/or more accurate processing by using modern computing devices.
Claims 4 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Wang-1 in view of Seo as applied to Claim 3, in further view of Zhang (CN 111797214 A).
Regarding Claim 4, Wang-1 in view of Seo teaches The asset creation method of claim 3.
Wang-1 in view of Seo does not explicitly disclose; however, Zhang teaches wherein the similarity is measured to correspond to a weighted sum of Jaccard similarity and cosine similarity (
“. . . the calculated Jaccard similarity, BM25 similarity, cosine similarity and edit distance similarity according to each weight value to weighted sum to obtain the similarity value between the question sentence and the query result; filtering the inquiry result of the similarity value not in the preset range.” Zhang p. 10.).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Zhang’s weighted sum of similarity values with Wang-1 in view of Seo. One of ordinary skill in the art would be motivated to provide greater flexibility to handle different scenarios by adjusting the weights for different similarity measures.
Claim 12 is substantially similar to Claim 4. The rejections analyses based on Wang-1 in view of Seo and Zhang for Claim 4 are applied to Claim 12.
Claims 5 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Wang-1 in view of Seo as applied to Claim 3, in further view of Chien et al. (US 20230386062 A1)
Regarding Claim 5, Wang-1 in view of Seo teaches The asset creation method of claim 3, wherein the similarity is measured based on embedding of position information and embedding of color information (
Wang-1 p. 8 right col.:
PNG
media_image5.png
274
782
media_image5.png
Greyscale
.
Here, the color image has embedded color information and position information, indicating where the color pixels are located. In addition, depth image also contains position information.
Measured similarity as explained in Claims 1-3 is based on the input images.).
Wang-1 in view of Seo does not explicitly disclose; however, Chien teaches wherein the similarity is measured based on embedding of position information and embedding of color information
PNG
media_image6.png
112
866
media_image6.png
Greyscale
“It can be understood that the calculation of loss value of the depth estimation model combines the mean square error and cosine similarity, which can not only improve the prediction accuracy of the depth estimation model, but also improve the color-sensitivity of the depth estimation model, allowing color differences between individual pixels even in low-texture areas to be distinguished. Obtaining depth images by using the depth estimation model can improve the accuracy of depth information.” Chien ¶ 64.).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Chien’s prediction method with Wang-1 in view of Seo. One of ordinary skill in the art would be motivated to enhance accuracy of predictions. “It can be understood that the calculation of loss value of the depth estimation model combines the mean square error and cosine similarity, which can not only improve the prediction accuracy of the depth estimation model, but also improve the color-sensitivity of the depth estimation model, allowing color differences between individual pixels even in low-texture areas to be distinguished. Obtaining depth images by using the depth estimation model can improve the accuracy of depth information.” Chien ¶ 64.
Claim 13 is substantially similar to Claim 5. The rejections analyses based on Wang-1 in view of Seo and Chien for Claim 5 are applied to Claim 13.
Claims 6 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Wang-1 in view of Seo as applied to Claim 3, in further view of Ezhov et al. (US 20210217170 A1).
Regarding Claim 6, Wang-1 in view of Seo teaches The asset creation method of claim 3, wherein creating the asset further comprises: performing
“In MMLNet, the encoder part uses DenseNet architecture that contains four modules. Each module is composed of a Dense Block, a convolution layer and an average pooling layer with stride 2.”
Wang-1 p. 7 right col.:
PNG
media_image7.png
432
772
media_image7.png
Greyscale
PNG
media_image8.png
168
568
media_image8.png
Greyscale
).
Wang-1 in view of Seo does not explicitly disclose; however, Ezhov teaches the DenseNet could be 3D convolution (“Model: The classification model has a DenseNet architecture. The only difference between the original and implementation of DenseNet by the present invention is a replacement of the 2D convolution layers with 3D ones.” Ezhov ¶ 63.).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Ezhov’s 3D convolution with Wang-1 in view of Seo. One of ordinary skill in the art would be motivated to conduct or enhance image classification. “Model: The classification model has a DenseNet architecture. The only difference between the original and implementation of DenseNet by the present invention is a replacement of the 2D convolution layers with 3D ones.” Ezhov ¶ 63.
Claim 14 is substantially similar to Claim 6. The rejections analyses based on Wang-1 in view of Seo and Ezhov for Claim 6 are applied to Claim 14.
Claims 8 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Wang-1 in view of Seo as applied to Claim 7, in further view of Li et al. (CN 116738208 A).
Regarding Claim 8, Wang-1 in view of Seo teaches The asset creation method of claim 7.
Wang-1 in view of Seo does not explicitly disclose; however, Li teaches wherein performing the dimension reduction comprises: applying mean pooling of 3D convolution to each point (
Li teaches convolution to each point, stating “Preferably, the convolution module 520 is further configured to determine the abstract expression corresponding to each original feature point according to the following formula: . . .. In the formula, xjO represents a real part convolution result in the jth base state, and yjO represents an imaginary part convolution result in the jth base state.” Li pp. 8-9.
Li teaches 3D nature of the convolution, stating “Wherein, the original point cloud comprises a plurality of original characteristic points, each original characteristic point is represented by the vector composed of three-dimensional space coordinate, illumination intensity, colour vector and so on.” Li p. 5.
Li teaches applying mean pooling , stating “The processing module 530 is further configured to: performing geometric mean pool processing on the abstract expression corresponding to each original characteristic point to obtain the pool-processed point cloud abstract characteristic expression, performing characteristic dimension reduction processing on the pool-processed point cloud abstract characteristic expression to obtain the quantum state point cloud combination characteristic.” Li p. 7.).
Therefore, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine Li’s mean pooling with Wang-1 in view of Seo. One of ordinary skill in the art would be motivated to compress data representation to reduce data storage needs and/or to simplify computation.
Claim 16 is substantially similar to Claim 8. The rejections analyses based on Wang-1 in view of Seo and Li for Claim 8 are applied to Claim 16.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Cao et al. (“Attention-based video object segmentation algorithm”), which teaches segmentation and LSTM, two important features in the claimed invention:
“To improve the segmentation performance on videos with large motion or deformation, ASCNN-LSTMVOS is proposed by constructing two branches: NNoP and OFbP. In NNoP, the attention mechanism is employed to highlight the features of objects. Then, to well capture temporal information, Conv3D is adopted to obtain the short-term temporal features, and AR-ConvLSTM is designed to capture the long–short-term temporal features under the interference of redundant frames.”
However, Cao is not as strong as Wang-1. For example, there is no clear asset creation, and LSTM appears to be part of the segmentation.
Gubbi Lakshminarasimha et al. (US 20220019804 A1), which teaches segmentation and LSTM, two important features in the claimed invention:
[0042] In yet another embodiment, wherein an architecture is proposed for spatio-temporal video object segmentation (ST-VOS) network using the ResNet, I3D and the LSTM as shown in FIG. 11. Herein, the input to the ST-VOS network is a set of ten RGB frames of size (10, 224, 224, 3). The input is processed via a I3D model and an intermediate output of size (5, 28, 28, 64) is extracted. The output of the I3D model is given as input to an ST-LSTM sub-network, consisting of four layers with filter sizes as shown in FIG. 11. Further, the intermediate outputs are collected and then feed it to a 3D deconvolution sub-network consisting of 3D convolution layers and 3D transpose convolution layers. The video frames input is parallelly passed through the ResNet and the outputs at multiple levels are captured. This is input to the corresponding levels in the deconvolution block to compute the final object segmentation map.
However, Gubbi is not as strong as Wang-1. For example, ST-LSTM appears to be part of the segmentation.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ZHENGXI LIU whose telephone number is (571)270-7509. The examiner can normally be reached M-F 9 AM - 5 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at 571-272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ZHENGXI LIU/Primary Examiner, Art Unit 2611