Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) was filed on 05/30/2025. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Election/Restrictions
Applicant’s election without traverse of claims 11-15 and 18-19 in the reply filed on 07/15/2026 is acknowledged.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 11-15 and 18-19 are rejected under 35 U.S.C. 103 as being unpatentable over Long et.al. (Long, Shangbang, et al. "Towards end-to-end unified scene text detection and layout analysis." 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2022. (Year: 2022)) in view of Reisswig et. al. (US Pub. No. 20210150201 A1) and further in view of Liu et. al. (Liu, Yuliang, et al. "Abcnet v2: Adaptive bezier-curve network for real-time end-to-end text spotting." IEEE Transactions on Pattern Analysis and Machine Intelligence 44.11 (2021): 8048-8064. (Year: 2021)
As per claim 11, Long teaches “A method performed by one or more computers, the method comprising:
receiving an original image depicting one or more text instances and a plurality of input object queries comprising learnable positional… corresponding to different regions of the original image; and” (See page 5 fig. 3, it shows an input image with one or more text instances. See also page 4 section 3.1 Data collection that shows the input images. See page 4 section 4.2 Model Architecture paragraphs 1-3 “Our unified detector is based on the recent Max-DeepLab [53] end-to-end panoptic segmentation framework. In this framework, we augment the in put pixels with a set of N learned object queries that are D−dimensional. Then we feed the pixels and object queries into a transformer-based encoder, the MaX-DeepLab back bone, in which the bidirectional communication between pixels and object queries allows the model to encode text instances in each of the object queries. With the encoded queries and pixel features, the text detection branch produces the text mask output, {ˆmi}N i=1.” “Backbone: The MaX-DeepLab[53] backbone is composed of an alternating stack of hourglass [39] style CNNs and the proposed dual-path transformer. The Hourglass style [39] CNNs are applied to pixel features. They encode features from coarse to fine resolutions iteratively and thus can produce high resolution features. The dual-path transformer [53] allows bidirectional communication between pixel features and the learnable object queries. It enables attention within pixel space and interaction among object queries. This makes it possible to encode long-range information in pixel features, and allows object queries to locate and retrieve text objects exclusively from pixels.” The object queries can be learned and locate object from pixels (by locating there is a positional embedding.) See also page 6 fig. 4 which shows the located object masks comprising text on different regions, also seen in fig. 3. See also page 5 section 4.3 Training targets. Long )
“processing the original image and the plurality of input object queries using a unified detector… neural network, wherein the unified detector… neural network comprises:” (See page 4 section 4.1, it uses a unified detector for end-to-end training using neural networks. See page 5 section 4.3 “Unified detector enables end-to-end training for both the scene text detection task and the layout analysis task…”. See page 4 section 4.2 “The architecture of the proposed unified detector is illustrated in Fig. 3. Our unified detector is based on the recent Max-DeepLab [53] end-to-end panoptic segmentation framework. In this framework, we augment the in put pixels with a set of N learned object queries that are D−dimensional…Backbone: The MaX-DeepLab[53] backbone is composed of an alternating stack of hourglass [39] style CNNs and the proposed dual-path transformer.” See also page 3 section 3.1 paragraph 2.The masks can be interpreted as polygons. Long)
“a feature extractor neural network configured to generate a set of encoded object queries, wherein the encoded object queries comprise an embedding corresponding to each region in the original image, and a set of encoded image pixel features; and” and “encoded object queries” (See page 5 fig. 3, it shows the encoded object queries along with the pixel features. See also page 4 section 3.1 Data collection that shows the input images. See page 4 section 4.2 Model Architecture paragraphs 1-3 “Our unified detector is based on the recent Max-DeepLab [53] end-to-end panoptic segmentation framework. In this framework, we augment the input pixels with a set of N learned object queries that are D−dimensional. Then we feed the pixels and object queries into a transformer-based encoder, the MaX-DeepLab back bone, in which the bidirectional communication between pixels and object queries allows the model to encode text instances in each of the object queries. With the encoded queries and pixel features, the text detection branch produces the text mask output, {ˆmi}N i=1.” “Backbone: The MaX-DeepLab[53] backbone is composed of an alternating stack of hourglass [39] style CNNs and the proposed dual-path transformer. The Hourglass style [39] CNNs are applied to pixel features. They encode features from coarse to fine resolutions iteratively and thus can produce high resolution features. The dual-path transformer [53] allows bidirectional communication between pixel features and the learnable object queries. It enables attention within pixel space and interaction among object queries. This makes it possible to encode long-range information in pixel features, and allows object queries to locate and retrieve text objects exclusively from pixels.” The object queries can be learned and locate object from pixels (since the object queries are learnable and have the ability to locate the text objects with attention within a space, it means that it implicitly uses a positional embedding.) Examiner interprets “embedding” as an “encoding”/”attachment”/”insertion” that can locate objects in the respective object query as also. . See also page 6 fig. 4 which shows the located object masks comprising text on different regions, also seen in fig. 3. See also page 5 subsection Layout branch “Layout branch: Layout branch takes the encoded queries from the backbone as the sole input. In order to separate layout features from text detection features, we apply an extra projection head for cluster embedding projection. For this projection head, we adopt a 3-layered multi-head self-attention layer [51] to obtain the normalized layout features, denoted as h ∈ RN×C.”. See also page 5 section 4.3 Training targets. Long ), however Long does not teach “positional embeddings” and “a polygon detector prediction neural network configured to process the encoded object queries to detect a polygon text boundary defined by a plurality of control points around each of one or more text instances in the original image.”
Reisswig teaches “learnable positional embeddings” (See paragraphs 10-12 and 21-28. “[0010]… For example, if the label generation process uses a serialized machine learning or artificial intelligence format, the position parameters may be embedded with the symbols to preserve the positional information. Using these embeddings, labels may be generated using the positional information to achieve higher accuracy with an accelerated learning process. ” “[0028] Label system 110 may receive document 120 and generate positional embeddings and/or labels to identify values of document 120. Label system 110 may include position vector network 112, label processor 114, and/or label network 116 to process document 120. Label processor 114 may manage position vector network 112 and/or label network 116. Label processor 114 may include one or more processors and/or memory configured to implement neural networks or machine learning algorithms. Position vector network 112 may be a neural network and/or other machine learning model configured to identify positional embeddings of characters, words, symbols, and/or tokens of document 120. Label network 116 may use these positional embeddings and/or word vector values to generate a label. Label processor 114 may control this process.” “[0033] After identifying the positional embeddings, label processor 114 may combine the positional embeddings with the input word vectors as inputs to label network 116. For example, a positional embedding vector “xi” may be identified for each input word vector “wi”. These positional embedding vectors may map information about the location of a particular token of the two-dimensional document into a higher dimensional space. The dimension may be the same or different from the input word vectors. Label processor 114 may combine and/or append the positional embedding vectors to the input word vectors and supply the combination as inputs to label network 116. In this manner, the combination may preserve the two-dimensional layout information for the label generation process.” See also paragraphs 33-44. By preserving position information of text, it is learning the positional embeddings. Reisswig)
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings of Long with the teachings of Reisswig for encoded object queries contain learnable positional embeddings. The modification would have been motivated by the desire to have higher accuracy, accelerated learning, preserve location information, therefore it is an improvement, less resource intensive, as suggested by Reisswig (See paragraphs 10, [0010]… The embodiments disclosed herein may analyze a document image to identify a sequence of position parameters for symbols or tokens of the document image. These position parameters may be used to preserve the layout information of the tokens in the document image and may provide increased accuracy during label generation. For example, if the label generation process uses a serialized machine learning or artificial intelligence format, the position parameters may be embedded with the symbols to preserve the positional information. Using these embeddings, labels may be generated using the positional information to achieve higher accuracy with an accelerated learning process.” See also paragraph 25 “[0025] By combining the positional embeddings with the vector values of the words, the label system may preserve the positional information when generating labels. The label system may utilize this positional information in the second neural network when generating labels to generate more accurate results. This configuration may further preserve positional information for use even when the second neural network uses a one-dimensional and/or sequential formatting. For example, this configuration may extract data from a table and preserve the table organization. In this manner, the two-dimensional information of a document may be preserved and utilized even in one-dimensional language models or neural networks. This configuration may also be less computer resources intensive and may be more efficient in training a machine learning model. This process may accelerate the machine learning process and also yield higher accuracy. Additionally, the neural network configuration may use less layers allowing for less resource intensive processing. In this manner, the configuration may be light-weight and fast when classifying characters and/or words of a document while still capturing the positional embeddings of each character or word.” See also paragraph 34. Reisswig)
Liu teaches “…polygon neural network” and “a polygon detector prediction neural network configured to process… to detect a polygon text boundary defined by a plurality of control points around each of one or more text instances in the original image.” (See page 4 fig.2 and section 3 Our Method, “An intuitive pipeline of our method is shown in Fig. 2. Inspired by [78], [79], [80], we adopt a single-shot, anchor free convolutional neural network as the detection framework… Next, we present the key components of the proposed ABCNet v2 in six components…:” along with section 3.2 Bezier Curve Detection paragraphs 1-4 “To simplify the detection for arbitrarily-shaped scene text instances, we propose to fit a Bezier curve by regressing several key points. The Bezier curve represents a parametric curve c(t) that uses the Bernstein Polynomials as its basis. The definition is shown in Equation (1)… To fit arbitrary shapes of the text with Bezier curves, we examine the arbitrarily shaped scene text from the existing datasets and we empirically show that a cubic Bezier curve (i.e., n ¼ 3) is sufficient to fit different formats of curved scene text, especially on dataset with word-level annotation.” See also fig.3 and page 5 column 1 paragraphs 1-4 “Fig.3 Cubic Bezier curves bi represents the control points. The green lines form a control polygon, and the black curve is the cubic Bezier curve. Note that with only two end-points b1 and b4, the Bezier curve degenerates to a straight line.” “To learn the coordinates of the control points, we first generate the Bezier curve annotations described in Section 3.1.1 and follow a similar regression method as in [34] to regress the targets. For each text instance, we use…” See also pages 5-6 section 3.3 BezierAlign. Liu)
It would have been obvious to one of ordinary skill in the art before the effective filing
date of the claimed invention to combine the teachings of Long with the teachings of Reisswig and Liu for the unified detector neural network be based on polygon detection to detect polygon text boundaries defined by control points. The modification would have been motivated by the desire to improve detection performance, improve recognition accuracy when sampling features of text regions, improve precision, more simplicity, more efficient and improve inference time, therefore it is an improvement, as suggested by Liu (See page 6 fig. 5 “Fig. 5. Comparison between previous sampling methods and BezierAlign. The proposed BezierAlign can more accurately sample features of the text region, which is essential for achieving good recognition accuracy. Note that the alignment procedure is applied to intermediate convolution features” See page 2 fig. 1. See abstract “Here, we tackle end-to-end text spotting by presenting Adaptive Bezier Curve Network v2 (ABCNet v2). Our main contributions are four-fold: 1) For the first time, we adaptively fit arbitrarily-shaped text by a parameterized Bezier curve, which, compared with segmentation-based methods, can not only provide structured output but also controllable representation. 2) We design a novel BezierAlign layer for extracting accurate convolution features of a text instance of arbitrary shapes, significantly improving the precision of recognition over previous methods. 3) Different from previous methods, which often suffer from complex post-processing and sensitive hyper-parameters, our ABCNet v2 maintains a simple pipeline with the only post-processing non-maximum suppression (NMS). 4) As the performance of text recognition closely depends on feature alignment, ABCNet v2 further adopts a simple yet effective coordinate convolution to encode the position of the convolutional filters, which leads to a considerable improvement with negligible computation overhead. Comprehensive experiments conducted on various bilingual (English and Chinese) benchmark datasets demonstrate that ABCNet v2 can achieve state-of-the-art performance while maintaining very high efficiency. More importantly, as there is little work on quantization of text spotting models, we quantize our models to improve the inference time of the proposed ABCNet v2.”)
Claims 18 and 19 are rejected under the same analysis as claim 11. (Reiswig teaches computer and storage. See paragraphs 61-63. Reiswig) (See page 1 abstract, it shows a link for the github repository which to one of ordinary skill in the art knows that it can be downloaded to storage and used on a computer. Page 7 section 5.3 Main Result shows that it uses computational resources. Long)
As per claim 12, Long in view of Reisswig and Liu teaches “The method of claim 11, wherein the polygon detector prediction neural network comprises:
a location head that is configured to generate a plurality of coordinates defining an axis-aligned bounding box for each text instance in the original image space;” (See Liu, page 6 fig. 4 and fig. 5 show the resulting axis-aligned bounding box of text instances in an image. See page 5 column 1 paragraphs 1-4 “To learn the coordinates of the control points, we first generate the Bezier curve annotations described in Section 3.1.1 and follow a similar regression method as in [34] to regress the targets. For each text instance, we use…” See also page 5 section 3.2 shows the generated coordinates used for an axis aligned box for each text instance “As pointed out in [14], conventional convolutions show limitation when learning a mapping between coordinates in (x,y) Cartesian space and coordinates in one-hot pixel space. The problem can be effectively solved by concatenating the coordinates to the feature maps. The recent practice of encoding relative coordinates [15] also show that the relative coordinates can provide informative cues for instance segmentation. Let fouts denotes the features of different scales of FPN, and Oi,x and Oi,y represent the absolute x and y coordinates, respectively, from all the locations (i.e., the location where the filters are generated) for the ith level of FPN. All Oi,x and Oi,y consist of two feature maps fox and foy. We simply concatenate fox and foy to the last channel of fouts along the channel dimension. Therefore, new features fcoord with additional two channels are formed, which are subsequently input to three convolutional layers with kernel size, stride, and padding size setting to 3, 1, and 1, respectively. We find that using such simple coordinate convolutions can considerably improve the performance of scene text spotting.” See also section 3.3 BezierAlign on pages 5-6 “…Unlike RoIAlign, the shape of the sampling grid of BezierAlign is not rectangular. Instead, each column of the arbitrarily-shaped grid is orthogonal to the Bezier curve boundary of the text. The sampling points have equidistant interval in width and height, respectively, which are bilinear interpolated with respect to the coordinates… Then the points of upper Bezier curve boundary tp and lower Bezier curve boundary bp are calculated according to Equation (1)… With the position of op, we can easily apply bilinear interpolation to calculate the result.” See also fig. 10. Liu)
“a shape head that is configured to generate a plurality of control points defining a polygon within a local bounding box image space as defined by the axis-aligned bounding box; and” (See Liu, page 6 fig. 4 and fig. 5 it shows the resulting axis-aligned bounding box of text instances in an image with a plurality of control points. See page 4 section 3.1 Bezier Curve Detection along with fig. 3 “To simplify the detection for arbitrarily-shaped scene text instances, we propose to fit a Bezier curve by regressing several key points. The Bezier curve represents a parametric curve… that uses the Bernstein Polynomials as its basis…Based on the cubic Bezier curve, we can formulate the arbitrarily-shaped scene text detection into a regression problem similar to bounding box regression, but with eight control points in total.” See page 5 column 1 paragraphs 1-4 “To learn the coordinates of the control points, we first generate the Bezier curve annotations described in Section 3.1.1 and follow a similar regression method as in [34] to regress the targets. For each text instance, we use…” See also page 5 section 3.2 shows the generated coordinates used for an axis aligned box for each text instance. See also section 3.3 BezierAlign on pages 5-6 “By exploiting the parameterization nature of a structured Bezier curve bounding box, we propose BezierAlign for feature sampling/alignment, which may be viewed as a flexible extension of RoIAlign . Unlike RoIAlign, the shape of the sampling grid of BezierAlign is not rectangular. Instead, each column of the arbitrarily-shaped grid is orthogonal to the Bezier curve boundary of the text. The sampling points have equidistant interval in width and height, respectively, which are bilinear interpolated with respect to the coordinates.” “…Unlike RoIAlign, the shape of the sampling grid of BezierAlign is not rectangular. Instead, each column of the arbitrarily-shaped grid is orthogonal to the Bezier curve boundary of the text. The sampling points have equidistant interval in width and height, respectively, which are bilinear interpolated with respect to the coordinates… Then the points of upper Bezier curve boundary tp and lower Bezier curve boundary bp are calculated according to Equation (1)… With the position of op, we can easily apply bilinear interpolation to calculate the result.” See also fig. 10. Liu)
“wherein the polygon detector prediction neural network is configured to scale and translate the polygon defined by the control points within the local bounding box space back into the original image space using the coordinates of the axis-aligned bounding box.” (See Liu, page 4 fig. 2 “Fig. 2. The framework of the proposed ABCNet v2. We use cubic Bezier curves and BezierAlign to extract multi-scale curved sequence features using the Bezier curve detection results. We concatenate coordinate channels to encode the position coordinates in FPN output features before sending to BezierAlign. The overall framework is end-to-end trainable with high efficiency. Here purple dots represent the control points of the cubic Bezier curve.” See also sections 3.1-3.4. See page 5 column 2 paragraph 3 “Let fouts denotes the features of different scales of FPN, and Oi,x and Oi,y represent the absolute x and y coordinates, respectively, from all the locations (i.e., the location where the filters are generated) for the ith level of FPN. All Oi,x and Oi,y consist of two feature maps fox and foy. We simply concatenate fox and foy to the last channel of fouts along the channel dimension…” See section 3.3 BezierAlign “ By exploiting the parameterization nature of a structured Bezier curve bounding box, we propose BezierAlign for feature sampling/alignment, which may be viewed as a flexible extension of RoIAlign. Unlike RoIAlign, the shape of the sampling grid of BezierAlign is not rectangular. Instead, each column of the arbitrarily-shaped grid is orthogonal to the Bezier curve boundary of the text. The sampling points have equidistant interval in width and height, respectively, which are bilinear interpolated with respect to the coordinates… Then the points of upper Bezier curve boundary tp and lower Bezier curve boundary bp are calculated according to Equation (1). Using tp and bp, we can linearly index the sampling point op by Equation:” See fig. 5, the resulting aligned bounding box is put back into the image space already scaled and translated the coordinates by using interpolating from section 3.2 CoordConv to convert coordinates and properly fit the text instance. It is all motivated by the same reasons stipulated in the rejection of claim 11. Liu)
As per claim 13, Long in view of Reisswig and Liu teaches “The method of claim 12, wherein the shape head comprises a Bezier polygon prediction head that predicts 4(m+1) control points corresponding to two Bezier polylines of order m that characterize polylines that constitute the polygon text boundary.” (See Liu, pages 4-5 section 3.1 Bezier Curve Detection “Based on the cubic Bezier curve, we can formulate the arbitrarily-shaped scene text detection into a regression problem similar to bounding box regression, but with eight control points in total…To learn the coordinates of the control points, we first generate the Bezier curve annotations described in Section 3.1.1 and follow a similar regression method as in [34] to regress the targets where xmin and ymin represent the minimum x and y values of the 4 vertexes, respectively.… The advantage of predicting the relative distance is that it is irrelevant to whether the Bezier curve control points are beyond the image boundary. Inside the detection head, we only use one convolution layer with 4(n+1) (n is the number of Bezier Curve order) output channels to learn the Δx and Δy, which is nearly cost-free while the results can still be accurate. For each text instance, we use ” and sections 3.1.1, 3.2 and 3.3. See also fig. 3 “Cubic Bezier curves. bi represents the control points. The green lines form a control polygon, and the black curve is the cubic Bezier curve. Note that with only two end-points b1 and b4, the Bezier curve degenerates to a straight line.” In section 3.1.1 “Here m represents the number of the annotated points for a curved boundary. For Total-Text and SCUT-CTW1500, m is 5 and 7, respectively. t is calculated by using the ratio of the cumulative length to the perimeter of the poly-line. According to Equations (1) and (4), we convert the original poly-line annotation to a parameterized Bezier curve.” It then corresponds in two polylines as seen in section 3.3 BezierAlign with fig. 5 and fig. 4. Liu)
As per claim 14, Long in view of Reisswig and Liu teaches “The method of claim 11, further comprising processing the encoded object queries with the polygon detector prediction neural network, wherein processing the encoded object queries with the polygon detector prediction neural network further comprises:
processing the encoded object queries with a layout head to generate a plurality of layout feature masks that correspond with each encoded object query;” (See Long, page 5 column 1 fig. 3 and paragraphs 1-3 “Figure 3. Illustration of our approach. Our method produces mask outputs for the text detection task and an affinity matrix for layout analysis task in a unified way. The text detection branch produces an N ×H×W tensor that represents the N softly exclusive masks. The layout analysis branch produces an N × N affin ity matrix that models the pairwise relationship of the predicted masks. The green links in the top right suggest clustering of the text instances, while the red links indicate the opposite. The binary classification score produced by textness branch is used to filter out non-text objects from object queries.” “Layout branch: Layout branch takes the encoded queries from the backbone as the sole input. In order to separate layout features from text detection features, we apply an extra projection head for cluster embedding projection. For this projection head, we adopt a 3-layered multi-head self-attention layer [51] to obtain the normalized layout features, denoted as h ∈ RN×C. We apply inner product of the lay out features followed by a sigmoid function with temperature τ to get the affinity matrix…” See also page 4 column 1 para 1 “For the second task of layout analysis, we also frame it as an instance segmentation task by treating each text cluster, i.e. “paragraph”, as one object in stance, following previous works [62]. The ground-truths for text lines and paragraphs are defined as the union of pixel-level masks of the underlying word level polygons.” Long)
“processing the encoded object queries with a textness head to generate a plurality of classification scores denoting a probability of each generated feature mask associated with each encoded object query being a text instance;” (See Long, page 5 column 2 paragraph 1-2 “Textness branch: The textness branch applies another 2 layered fully connected layers and a sigmoid function to produce the binary classification scores {ˆyi}N i=1.” See page 5 column 1 fig. 3 and paragraphs 1-3 “Figure 3. Illustration of our approach. Our method produces mask outputs for the text detection task and an affinity matrix for layout analysis task in a unified way. The text detection branch produces an N ×H×W tensor that represents the N softly exclusive masks. The layout analysis branch produces an N × N affin ity matrix that models the pairwise relationship of the predicted masks. The green links in the top right suggest clustering of the text instances, while the red links indicate the opposite. The binary classification score produced by textness branch is used to filter out non-text objects from object queries.” See also section 4.3 Training targets. See also pages 4-5 subsection Text detection branch “Text detection branch: The text detection branch takes the outputs of the MaX-DeepLab backbone and produces the text mask outputs. Two fully-connected layers produce mask queries from the encoded queries, denoted as f ∈ RN×D. Similarly, two convolutional layers produce normalized pixel features, denoted as g ∈ RD×H×W. The text mask prediction is the inner product of f and g:” Long)
“computing an inner product of the generated feature masks from the layout head to produce an affinity matrix of feature similarity; and” (See Long, page 5 column 1 fig. 3 and paragraphs 1-3 “Figure 3. Illustration of our approach. Our method produces mask outputs for the text detection task and an affinity matrix for layout analysis task in a unified way. The text detection branch produces an N ×H×W tensor that represents the N softly exclusive masks. The layout analysis branch produces an N × N affinity matrix that models the pairwise relationship of the predicted masks. The green links in the top right suggest clustering of the text instances, while the red links indicate the opposite. The binary classification score produced by textless branch is used to filter out non-text objects from object queries.” “Layout branch: Layout branch takes the encoded queries from the backbone as the sole input. In order to separate layout features from text detection features, we apply an extra projection head for cluster embedding projection. For this projection head, we adopt a 3-layered multi-head self-attention layer [51] to obtain the normalized layout features, denoted as h ∈ RN×C. We apply inner product of the lay out features followed by a sigmoid function with temperature τ to get the affinity matrix…”. See also section 4.3 Training targets Long)
“generating a paragraph grouping representation for each encoded object query that is determined to be a text instance from the affinity matrix.” (See Long, page 5 column 1 fig. 3 and paragraphs 1-3 “Figure 3. Illustration of our approach. Our method produces mask outputs for the text detection task and an affinity matrix for layout analysis task in a unified way. The text detection branch produces an N ×H×W tensor that represents the N softly exclusive masks. The layout analysis branch produces an N × N affinity matrix that models the pairwise relationship of the predicted masks. The green links in the top right suggest clustering of the text instances, while the red links indicate the opposite. The binary classification score produced by textness branch is used to filter out non-text objects from object queries.” “Layout branch: Layout branch takes the encoded queries from the backbone as the sole input. In order to separate layout features from text detection features, we apply an extra projection head for cluster embedding projection. For this projection head, we adopt a 3-layered multi-head self-attention layer [51] to obtain the normalized layout features, denoted as h ∈ RN×C. We apply inner product of the lay out features followed by a sigmoid function with temperature τ to get the affinity matrix…”. See also section 4.3 Training targets. See also page 4 subsection Unified layout analysis: “Unified layout analysis: Unified detector analyzes the lay out and performs text clustering by producing an affinity matrix: ˆ A∈[0,1]N×N. Entry ˆAi,j in this matrix represents the probability of text represented by mi and mj belonging to the same semantic/paragraph group” See also page 4 column 1 para 1 “For the second task of layout anal ysis, we also frame it as an instance segmentation task by treating each text cluster, i.e. “paragraph”, as one object in stance, following previous works [62]. The ground-truths for text lines and paragraphs are defined as the union of pixel-level masks of the underlying word level polygons.” Long)
As per claim 15, Long in view of Reisswig and Liu teaches “The method of claim 14, wherein the unified detector polygon neural network has been trained to minimize a loss function that comprises a weighted sum of one or more losses, wherein the losses comprise, for each predicted text instance:
a text loss that characterizes an overlap of the predicted text mask from the textness head with a ground truth text mask in the original image space;
a paragraph layout analysis loss that characterizes whether the affinity matrix generated by the layout head maps each text line in the text instance to a correct paragraph;
an original image space polygon loss that characterizes the overlap of the polygon text boundary detected by the polygon detector prediction neural network with a respective ground truth polygon in the original image space;
an original image space location loss that characterizes an accuracy of the plurality of coordinates predicted by the location head that defines the axis-aligned bounding box in the original image space; and
a local axis-aligned bounding box space polygon loss that characterizes the accuracy of the plurality of polygon control points predicted by the shape head in the local axis-aligned bounding box space.” (Since the claim language mentions that it uses one or more of the weight losses, within a BRI (Broadest Reasonable Interpretation), it only needs at least one of the losses for the claim to be rejected. See Liu, pages 5-6, section 4.3 Training targets along with subsections Text detection loss and layout analysis loss along with equations 7-10. On page 5 column 2 to page 6 column 1 “Unified detector enables end-to-end training for both the scene text detection task and the layout analysis task. The key ingredient is to perform a bipartite matching between prediction and groundtruth since our model produces an unordered set of outputs. We first describe the matching between prediction and groundtruth of the detection task and the metric we use. Then we show the joint optimization of our unified detector for both tasks. For a pair of prediction ( ˆmi, ˆyi) and groundtruth (mj,yj), the score is defined as: sim(i, j) = [ˆyiyj +(1− ˆyi)(1−yj)]×Dice(ˆmi,mj) (5) where Dice(ˆmi,mj) denotes the Dice coefficient [36] between the pair of masks. It measures mask similarity. This score considers both the classification score and mask score. The goal of matching is to find a permutation of N el ements σ ∈ GN to maximize the total similarity between predictions and groundtruths:” “Text detection loss: The training target for text detection is adopted from MaX-DeepLab … where dotted variables y^¨i and Di¨ce(m^i, mσ(i)) denote constant weights and gradients do not pass through them. α is a balancing factor between positive and negative samples. Layout analysis loss: We first define the ground-truths for output of layout analysis branch. Each text instance comes with a text cluster ID, denoted as {ci}N i=1. This is part of the annotations of the proposed HierText dataset. The groundtruth affinity matrix can be intuitively defined as: A[i, j] = 1(ci == cj) Then, the layout analysis loss can be computed as… The final training target is the weighted sum of the text detection loss Ldet, the layout analysis loss Llay. We also find it useful to incorporate the semantic segmentation loss Lseg and instance discrimination loss Lins as defined in MaX-DeepLab [53]. As a result, the model is jointly optimized for the following loss function…” See also Table 4. See also page 4 column 1 para 1 “For the second task of layout anal ysis, we also frame it as an instance segmentation task by treating each text cluster, i.e. “paragraph”, as one object in stance, following previous works [62]. The ground-truths for text lines and paragraphs are defined as the union of pixel-level masks of the underlying word level polygons.” Liu)
Pertinent Prior Art
Xu, Yang, et al. "Layoutlmv2: Multi-modal pre-training for visually-rich document understanding." Proceedings of the 59th annual meeting of the association for computational linguistics and the 11th international joint conference on natural language processing (volume 1: Long papers). 2021. (Year: 2021), also teaches axis-aligned bounding boxes and positional embeddings.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DYLAN J MENDEZ MUNIZ whose telephone number is (703)756-5672. The examiner can normally be reached M-F, 8AM - 5PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vu Le can be reached at (571) 272-7332. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DYLAN JOHN MENDEZ MUNIZ/Examiner, Art Unit 2675
/VU LE/Supervisory Patent Examiner, Art Unit 2668