DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
The claims will be read under the broadest reasonable interpretation standard outlined in
MPEP § 2111.01.
Specification
In paragraphs [0009], [0014], and [0052] “one more images” reads as a typographical error for “one or more images”
In paragraph [0036], “one more image capture devices” reads as a typographical error for “one or more image capture devices”
In paragraph [0064], “one more non-transitory” reads as a typographical error for “one or more non-transitory”
In paragraph [0010], “second distortion correction orientation may even be further based on” reads as a typographical error for “second distortion correction operation may even be further based on”
In paragraph [0030], “the initially-defined second portion may then tracked” reads as a grammatical error for “the initially-defined second portion may then be tracked”
Appropriate correction is required.
Claim Objections
Claim 1 is objected to for the following informalities: “one more images” reads as a typographical error for “one or more images”.
Claim 4 is objected to for the following informalities: “the first output image distortion-corrected first portion” lacks direct antecedent basis as only “the distortion-corrected first portion” was previously recited.
Claim 6 is objected to for the following informalities: “connected to first” reads as a grammatical error for “corrected to the first”; “via being embedded in first electronic device” reads as a grammatical error for “via being embedded in the first electronic device”
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because they are directed to ineligible patent subject matter. The claims are directed to the Abstract Idea grouping of mental processes under MPEP § 2106.04(a)(2)(III) and mathematical calculations under MPEP § 2106.04(a)(2)(I). These are judicial exceptions under Step 2A, Prong One of the framework established by the cases of Alice Corp. Pty. Ltd. v. CLS Bank Int'l, 573 U.S. 208, 216, 110 USPQ2d 1976, 1980 (2014) and Mayo Collaborative Servs. v. Prometheus Labs., Inc., 566 U.S. 66, 71, 101 USPQ2d 1961, 1965 (2012). See MPEP § 2106.04(II).
PNG
media_image1.png
200
400
media_image1.png
Greyscale
Step 1: The claims in question are directed to a method, device, and non-transitory computer-readable media (CRM) for “simultaneous subject and desk capture during videoconferencing.” Machines (devices), processes (methods), and articles of manufacture (non-transitory CRM) are all statutory categories. See MPEP 2106.03(I), “A machine is a "concrete thing, consisting of parts, or of certain devices and combination of devices." Digitech, 758 F.3d at 1348-49, 111 USPQ2d at 1719 (quoting Burr v. Duryee, 68 U.S. 531, 570, 17 L. Ed. 650, 657 (1863)). This category "includes every mechanical device or combination of mechanical powers and devices to perform some function and produce a certain effect or result." Nuijten, 500 F.3d at 1355, 84 USPQ2d at 1501 (quoting Corning v. Burden, 56 U.S. 252, 267, 14 L. Ed. 683, 690 (1854))”; See MPEP 2106.03(I), “NTP, Inc. v. Research in Motion, Ltd., 418 F.3d 1282, 1316, 75 USPQ2d 1763, 1791 (Fed. Cir. 2005) ("[A] process is a series of acts.") (quoting Minton v. Natl. Ass’n. of Securities Dealers, 336 F.3d 1373, 1378, 67 USPQ2d 1614, 1681 (Fed. Cir. 2003)). As defined in 35 U.S.C. 100(b), the term "process" is synonymous with "method."”; See MPEP 2106.03(I), “A manufacture is "a tangible article that is given a new form, quality, property, or combination through man-made or artificial means." Digitech, 758 F.3d at 1349, 111 USPQ2d at 1719-20 (citing Diamond v. Chakrabarty, 447 U.S. 303, 308, 206 USPQ 193, 197 (1980)). As the courts have explained, manufactures are articles that result from the process of manufacturing, i.e., they were produced "from raw or prepared materials by giving to these materials new forms, qualities, properties, or combinations, whether by hand-labor or by machinery." Samsung Electronics Co. v. Apple Inc., 137 S. Ct. 429, 120 USPQ2d 1749, 1752-3 (2016) (quoting Diamond v. Chakrabarty, 447 U. S. 303, 308, 206 USPQ 193, 196-97 (1980)); Nuijten, 500 F.3d at 1356-57, 84 USPQ2d at 1502.”; See MPEP 2016.03(II) (distinguishing statutory random-access memory and non-statutory carrier waves, as relevant to claim 20). (Step 1: Yes).
Step 2A, Prong One: As explained in MPEP 2106.04(II), a claim “recites” a judicial
exception when the judicial exception is “set forth” or “described” in the claim. Here, each
claim recites or depends upon the mental processes of obtaining (an image), cropping (a portion of the image), and transmitting an image (Claim 1, “obtaining…one or more images…cropping a first portion of a first image…transmitting the…portion…”; Claim 5 “obtaining orientation information…”; Claim 12, “the portion of the surface is defined by a user gesture”; Claim 18, “an occlusion detection operation”). The claims further recite or depend upon various mathematical calculations (distortion correction, image rotation, compositing, as functions acting upon a matrix; See MPEP § 2106.04(d), “simply implementing a mathematical principle on a physical machine, namely a computer, was not a patentable application of that principle”).
The claims are recited at a high level of generality and lack any specifics precluding such
an analysis from being interpreted under the mental processes grouping of “practically performed
in the mind” (see also MPEP § 2106.04(a)(2) identifying how e.g. a use of pen and paper, a ruler,
or a computer as a tool (to assist in visually/mentally analyzing/observing acquired
images/video) fails to preclude such an interpretation under the mental processes judicial
exception). Activities such as “at a first electronic device” or “a processor configured to” therefore may be performed mentally, even if they may require the additional computer tool. Similarly, high-level image obtaining/transfer/computer calculation does not elevate these claims past a mental process and/or mathematical calculation.
Regarding artificial intelligence, to the extent it is implicated, claim 9’s “deep neural network” is comparable to Claim 2 of Example 47 of the July 2024 PEG regarding subject matter eligibility (https://www.uspto.gov/sites/default/files/documents/2024-AISMEUpdateExamples47-49.pdf). As stated therein, an artificial intelligence’s analyses, detections, and reinforcement learnings may be practically performed in the human mind. To the extent mathematical calculations are required to operate and train the artificial intelligence in image analysis, the separate judicial exception is also implicated.
As such, the usage of a computer to obtain, edit, and transmit images at a high-level of generality does not elevate these claims beyond a mental process. (Step 2A, Prong One: Yes).
Step 2A, Prong Two: If Prong One of Step 2A is met, the examiner must consider (1)
whether there are any ‘additional elements’ recited in the claim beyond the judicial exception,
and (2) evaluate those additional elements individually and in combination to determine whether
the claim as a whole integrates the exception into a practical application. See MPEP §
2106.04(d).
Limitations the courts have found indicative of integration include: an improvement in
the functioning of a computer, or an improvement to other technology or technical field, as
discussed in MPEP §§ 2106.04(d)(1) and 2106.05(a); applying or using a judicial exception to
effect a particular treatment or prophylaxis for a disease or medical condition, as discussed in
MPEP § 2106.04(d)(2); implementing a judicial exception with, or using a judicial exception in
conjunction with, a particular machine or manufacture that is integral to the claim, as discussed
in MPEP § 2106.05(b); effecting a transformation or reduction of a particular article to a
different state or thing, as discussed in MPEP § 2106.05(c); and applying or using the judicial
exception in some other meaningful way beyond generally linking the use of the judicial
exception to a particular technological environment, such that the claim as a whole is more than
a drafting effort designed to monopolize the exception, as discussed in MPEP § 2106.05(e).
Limitations that the courts have found non-indicative of integration include: merely
reciting the words "apply it" (or an equivalent) with the judicial exception, or merely including
instructions to implement an abstract idea on a computer, or merely using a computer as a tool to
perform an abstract idea, as discussed in MPEP § 2106.05(f); adding insignificant extra-solution
activity to the judicial exception, as discussed in MPEP § 2106.05(g); and generally linking the
use of a judicial exception to a particular technological environment or field of use, as discussed
in MPEP § 2106.05(h).
As an additional note, ‘additional elements’ are generally limitations excluded from
interpretation under the Abstract Idea groupings, and may comprise portions of limitations
otherwise identified as falling under those Abstract Idea groupings of the 2019 PEG (e.g. any
‘determination’ that may be made mentally by a user, neural network and/or generic computer hardware is considered under the ‘apply it’ considerations of 2106.05(f)). Any ‘providing’/outputting broadly, and ‘collection/input’ of data (i.e obtaining of the image, obtaining of image information, transmitting of the final image), also fail(s) to integrate at least in view of MPEP 2106.05(g) (extra-solution data gathering/output) and/or 2106.05(h) as ‘generally linking’ the exception to a field of use involving machine learning and/or imagery so acquired (e.g. the use of a computer to acquire/transmit said image broadly). The same determination holds for dependent claims that serve to limit the collection/output of data/images (by means of what is collected based on recited conditions) and/or introduce limitations generally linking to a field of use.
None of the instant claims appear to explicitly/clearly capture/recite any disclosed
improvement in technology (see MPEP 2106.05(a), with note that ‘functioning of a computer’
concerns functions integral to the way a computer operates and not ‘functions’ that a generic
computer can be programmed/adapted to perform (see also 2106.05(f))) and any ‘additional
elements’, even when considered in combination, fail to integrate at Prong Two of Step 2A
accordingly. Integration in view of subsection (a) requires an identification of the manner in
which the improvement is achieved, to be explicitly and specifically recited in the claims, as
‘additional elements’ precluded from interpretation under any of the Abstract Idea groupings
(since the improvement cannot be to the exception itself). With reference to MPEP 2106.05(a):
It is important to note, the judicial exception alone cannot provide the improvement. The improvement can be provided by one or more additional elements. See the discussion of Diamond v. Diehr, 450 U.S. 175, 187 and 191-92, 209 USPQ 1, 10 (1981))
As applicable here, additional limitations not directed to a judicial exception fail to
integrate at Prong Two of Step 2A. Claim 1 recites a “a first electronic device, one or more images captured by a first image capture device”; Claim 2 recites a “video stream”; Claim 3 recites a “second electronic device”; Claim 6 recites a “connection”; Claim 8 recites “a detector”; Claim 9 recites a “deep neural network”; Claim 19 recites a “memory…processor…”; Claim 20 recites a “non-transitory computer readable medium”. The incorporation of broadly recited and conventional computer hardware, image capture and machine-learning systems does little more than generally link the judicial exceptions of mental processes and mathematical calculations to a field-of-use and technological environment. See MPEP §§ 2106.05(h); 2106.05(f).
Claim 1 recites “transmitting the distortion-corrected first portion and the distortion-corrected second portion to a second electronic device”. This limitation constitutes insignificant extra-solution activity under MPEP § 2106.05(g). Specifically, the limitations amount to no more than necessary data outputting, under rationale 3 of MPEP § 2106.05(g).
Even when viewed in combination, any additional elements present do not integrate the
recited judicial exception into a practical application (Step 2A, Prong Two: No), and the claims
are directed to the judicial exception. (Revised Step 2A: Yes → Step 2B).
Step 2B: If Prong Two of Step 2A is not met, the examiner must consider whether the
claim as a whole amounts to ‘significantly more’ than the recited exception, i.e., whether any
‘additional element’, or combination of additional elements, adds an inventive concept to the
claim. The considerations of Step 2A Prong 2 and Step 2B overlap, but differ in that 2B also
requires considering whether the claims feature any “specific limitation(s) other than what is
well-understood, routine, conventional activity in the field” (WURC) (MPEP § 2106.05(d)).
Such a limitation if specifically recited however, must still be excluded from interpretation under
any of the Abstract Idea groupings. Step 2B further requires a re-evaluation of any additional
elements drawn to extra-solution activity in Step 2A (e.g. obtaining the image, transmitting output) – however no limitations appear directed to any novel collection or transmitting per se. Limitations not indicative of an inventive concept/‘significantly more’ include those that are not specifically recited (instead recited at a high level of generality), those that are established as WURC (a plurality of cited references serve to evidence the WURC nature of ‘analysis’ based at least in part on corroborating/additional ground data), and/or those that are not ‘additional elements’ by nature of their analysis at Prong One of Step 2A (i.e. directed to the exception – see above re. deciding that a second acquisition may be advantageous/desired). The July 2024 PEG describes that an improvement/ inventive concept (for ‘significantly more’ determination(s)) cannot be to the judicial exception itself. As additionally recited by [0029] of the claimed invention’s specification and relevant to claim 9, the deep neural network is understood to include a plurality of pre-trained and conventional DNNs, to perform a function which can also be achieved by other broadly recited “detectors”.
The claims in question recite little beyond those limitations recited at a high level
of generality and falling under e.g. the mental process and mathematical calculation Abstract Idea groupings, and would monopolize the exceptions accordingly. The additional limitations of computer processing, machine-learning, and input/output as recited are WURC, as evidenced by the body of prior art cited by the examiner in this office action. (Step 2B: No).
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-2, 4-7, 10, 12-15, 17, and 19-20 are rejected under 35 U.S.C. § 102(a)(1) and 35 U.S.C. § 102(a)(2) as being anticipated by Kasatani et. al (US 20150109401 A1) (Hereinafter, “Kasatani”)
With respect to claim 1, Kasatani discloses:
A method ([0001]), comprising:
obtaining, at a first electronic device, one or more images captured by a first image capture device ([0009]; Figs. 6-12)
cropping a first portion of a first image of the one or more images, wherein the first portion comprises a face of a human subject in the first image ([0009]; [0020] “FIG. 10 is a view for explaining an image obtained by performing distortion correction processing on almost the center portion, right side portion, and left side portion of the image illustrated in FIG. 6 in the video-conference terminal device for the video-conference system according to the embodiment and combining them into an image”; [0097]-[0100]; [0174]-[0176]; Figs. 6-12)
applying a first distortion correction operation to the first portion ([0009]; [0059]- [0060]; [0096]-[0114] “As another area, “center, right and left” means that the distortion correction processing is performed on the clipped images obtained by cutting out almost the center portion, almost the right side portion, and almost the left side portion of the image, on which the distortion correction processing has not yet been performed, and the processed images are displayed together on one screen.”; [0169]; [0178]-[0179]; [0256]-[0259]; Figs. 6-12)
cropping a second portion of the first image of the one or more images, wherein the second portion comprises a portion of a surface in the first image ([0009]; [0095]-[0100]; [0164]; [0175] “FIG. 9 illustrates an example of distortion correction obtained by clipping the left side portion and the right side portion of the image exemplified in FIG. 6, performing the distortion correction processing on the clipped images, and displaying them together on one screen. This provides a clear vision of the facial expression of the two attendees on the sides in FIG. 6, for example, because they are also displayed in a large size”; [0196]; Figs. 6-12)
applying a second distortion correction operation to the second portion ([0009]; [0096]-[0114]; [0126]-[0127] “When the predetermined condition, is selected, the correction table is referred to for narrowing down the area on which the distortion correction processing is performed to specific ones…The distortion correction table is selected accordingly when the area on which the distortion correction processing is performed is selected”; [0133]-[0138]; [0169]; [0175]; [0178]-[0181] “The CPU switches the image correction table to the one appropriate for the distortion correction processing on the clipped area to select an intended area from the clipped area at Step S2…The CPU applies the image correction table selected at Step S3 to the selected intended area and performs the image distortion correction-processing on the area (Step S4).”; [0278]-[0293] discussing user-specified masking regions, “Another distortion correction processing method will be described in which a masked area can be readily specified for clipping or masking an intended area…The operator can specify a plurality of areas…”; Fig. 3; Figs. 6-12
The examiner notes that the claim under BRI does not require that a surface-specific distortion correction be performed in this second cropped portion; only that the second portion includes a portion of a surface. Accordingly, as Figs. 6-9 show that surfaces (such as of the table, or paper in the background) may be included (and even visibly corrected with respect to Fig. 7, Figs. 9-10) in these second cropped portions, the claim is met. The “surface” of claim 1 may also be considered as yet another human.)
PNG
media_image2.png
355
709
media_image2.png
Greyscale
transmitting the distortion-corrected first portion and the distortion-corrected second portion to a second electronic device ([0047]; [0054]; [0062]; [0143]; Fig. 2)
With respect to claim 2, Kasatani discloses:
The method of claim 1, wherein the one or more images comprises a video image stream ([0009])
With respect to claim 4, Kasatani discloses:
The method of claim 1, wherein the transmitting of the first output image distortion-corrected first portion and the distortion-corrected second portion to a second electronic device occurs prior to the obtaining of a second image of the one or more images, wherein the second image is captured subsequently to the first image ([0054]-[0055]; [0062] “The digital video signals of the selected image are transmitted to the video-conference terminal devices 40 through the network control unit 16 as an example of a second control unit and the network 30, and then displayed on the liquid crystal display unit 101 of the video-conference terminal devices 40. This enables the attendee at the conference on the site where the remote video-conference terminal device 40 is used to progress the conference while watching the image on which the distortion correction processing has been appropriately performed”)
With respect to claim 5, Kasatani discloses:
The method of claim 1, further comprising:
obtaining orientation information associated with the first image capture device during the capture of each of the one or more images ([0130]-[0131] “The “LCD slant” includes four detection stages of the LCD hinge movement 114: “90° or larger”, “75° to under 90°”, “60° to under 75°”, and “under 60°” based on the main body 100 a placed within the horizontal plane. Although the “LCD slant” column illustrated in FIG. 5 includes the four detection stages of the LCD hinge movement 114: “90° or larger”, “75° to under 90°”, “60° to under 75°”, and “under 60°”, the four detection stages may be represented with the values 0 to 3.”; [0201] “FIG. 12 is an exemplary flowchart relating to an image distortion correction method including a step of detecting the angle of elevation or the direction of the camera. The flowchart illustrated in FIG. 12 has the same steps as the flowchart illustrated in FIG. 11 except that Step S12 is included for detecting the angle of elevation or the direction of the wide-angle camera”; Fig. 11; Fig. 12)
With respect to claim 6, Kasatani discloses:
The method of claim 1, wherein the first image capture device is connected to a first electronic device in one of the following ways: a wired connection, a wireless connection, or via being embedded in first electronic device (Fig. 1; Fig. 2)
With respect to claim 7, Kasatani discloses:
The method of claim 1, wherein cropping a first portion of a first image of the one or more images further comprises cropping the first portion according to one or more predetermined framing rules ([0180] “The CPU then displays the image of the area corrected at Step S4 (Step S5). When the step of preliminarily clipping the area for displaying the image as described above, four dots for determining a rectangular area to be cut out is defined in advance in the distortion correction table. The distortion correction table is applied to only the rectangular area. This process is useful for using the data in the area outside of the clipped image”; [0278]-[0293])
With respect to claim 10, Kasatani discloses:
The method of claim 1, further comprising:
tracking a location and size of the portion of the surface across the one or more images captured by the first image capture device ([0278]-[0293]; [0284] “After that, when the operator specifies an area to be clipped or an area to be masked on his own image displayed on the (touch panel) liquid crystal display unit 101 with his finger or the tip of the stylus, the CPU forms a masking pattern according to a signal from the touch panel module (Step S42). The operator can specify a plurality of areas”; Fig. 20)
PNG
media_image3.png
935
1047
media_image3.png
Greyscale
With respect to claim 12, Kasatani discloses:
The method of claim 1, wherein the portion of the surface is defined by a user gesture captured at the first image capture device (Fig. 1; Fig. 20; [0278]-[0293]; [0284] “After that, when the operator specifies an area to be clipped or an area to be masked on his own image displayed on the (touch panel) liquid crystal display unit 101 with his finger or the tip of the stylus, the CPU forms a masking pattern according to a signal from the touch panel module (Step S42). The operator can specify a plurality of areas”; As the video display unit is directly interacted with, the camera of the device of Fig. 1 will necessarily observe motions associated with the gesture in a user-facing video stream)
With respect to claim 13, Kasatani discloses:
The method of claim 1, wherein the first and second distortion correction operations are different ([0096]-[0114]; [0126]-[0127]; [0133]-[0138]; [0169]; [0175]; [0178]-[0181] “The CPU switches the image correction table to the one appropriate for the distortion correction processing on the clipped area to select an intended area from the clipped area at Step S2…The CPU applies the image correction table selected at Step S3 to the selected intended area and performs the image distortion correction-processing on the area (Step S4).”; Fig. 3; Figs. 6-11)
With respect to claim 14, Kasatani discloses:
The method of claim 5, wherein the second distortion correction operation comprises applying a rotation operation to the second portion, based, at least in part, on orientation information associated with the first image capture device during the capture of the first image ([0222] “In the present embodiment, therefore, the image correction tables corresponding to the degrees of the LCD slant are prepared in advance. When the user operates the LCD hinge mechanism to change the degrees of the LCD slant, the image correction table corresponding to the degrees of the LCD slant is automatically selected.”; Fig. 5)
With respect to claim 15, Kasatani discloses:
The method of claim 14, wherein the orientation information associated with the first image capture device during the capture of the first image comprises one or more of pitch ([0037]; [0220]-[0221]; Fig. 1)
With respect to claim 17, Kasatani discloses:
The method of claim 1, further comprising:
compositing the distortion-corrected first portion and the distortion-corrected second portion into a first output image ([0019]-[0020] “Fig. 10 is a view for explaining an image obtained by performing distortion correction processing on almost the center portion, right side portion, and left side portion of the image illustrated in Fig. 6 in the video-conference terminal device for the video-conference system according to the embodiment and combining them into an image”; [0099]; [0274] “The “center, right and left” in the “clipped image” column in FIG. 18 represents the image in which the clipped areas, 1, 2, and 3 are combined together. For example, the “center, right and left” image corresponds to the image illustrated in FIG. 10”)
wherein transmitting the distortion-corrected first portion and the distortion-corrected second portion to the second electronic device comprises transmitting the first output image to the second electronic device ([0047] “Video content displayed on the liquid crystal display unit 101 may also be therefore displayed on an external monitor or an external projector”; [0054]; [0062]; [0143]; [0176] “FIG. 10 illustrates an image obtained by performing distortion correction processing on the respective images in FIGS. 8 and 9, unifying the size of the images including the respective attendees into almost the same size, and displaying them together on one screen. This provides a clear vision of the facial expression of all of the attendees, for example”)
With respect to claim 19, Kasatani discloses:
An electronic device ([Abstract]), comprising
a memory ([0044])
a first image capture device (Fig. 1)
a first positional sensor ([0076])
one or more processors operatively coupled to the memory, wherein the one or more processors are configured to execute instructions causing the one or more processors to ([0044]; [0170]):
obtain one or more images captured by the first image capture device ([0009])
crop a first portion of a first image of the one or more images, wherein the first portion comprises a face of a human subject in the first image ([0009]; [0020];[0097]-[0100]; [0174]-[0176]; Figs. 6-12)
apply a first distortion correction operation to the first portion ([0009]; [0059]- [0060]; [0096]-[0114]; [0169]; [0178]-[0179]; [0256]-[0259]; Figs. 6-12)
crop a second portion of the first image of the one or more images, wherein the second portion comprises a portion of a surface in the first image ([0009]; [0095]-[0100]; [0164]; [0175]; [0196]; Figs. 6-12)
apply a second distortion correction operation to the second portion, wherein the second distortion correction operation comprises applying a rotation operation to the second portion, based, at least in part, on orientation information obtained from the first positional sensor and associated with the capture of the first image ([0009]; [0037]; [0076]; [0096]-[0114]; [0126]-[0127]; [0133]-[0138]; [0169]; [0175]; [0178]-[0181]; [0278]-[0293]; [0220]-[0222]; Fig. 1; Fig. 3; Figs. 5-12)
transmit the distortion-corrected first portion and the distortion-corrected second portion to another electronic device ([0047]; [0054]; [0062]; [0143]; Fig. 2)
With respect to claim 20, Kasatani discloses:
A non-transitory computer readable medium comprising computer readable instructions executable by one or more processors ([0044]; [0170]) to:
obtain a first sequence of images captured by a first image capture device ([0009])
crop a first portion of a first image of the first sequence of images, wherein the first portion comprises a face of a human subject in the first image ([0009]; [0020]; [0097]-[0100]; [0174]-[0176]; Figs. 6-12)
apply a first distortion correction operation to the first portion ([0009]; [0059]-[0060]; [0096]-[0114]; [0169]; [0178]-[0179]; [0256]-[0259]; Figs. 6-12)
obtain a second sequence of images captured by a second image capture device ([0009]; [0054]-[0055]; [0053]; [0062]; [0143] “Once the video-conference server 50 starts the video-conference room, the video-conference server 50 receives the compressed digital video signals and the compressed digital audio signals from all of the video-conference terminal devices 100 and 40 attending in the video-conference room”; Fig. 1)
crop a second portion of a second image of the second sequence of images, wherein the second portion comprises a portion of a surface in the second image, and wherein the second image corresponds in time to the first image ([0009]; [0095]-[0100]; [0143]; [0164]; [0175]; [0196]; Fig. 1; Figs. 6-12)
apply a second distortion correction operation to the second portion, wherein the second distortion correction operation comprises applying a rotation operation to the second portion ([0009]; [0037]; [0076]; [0096]-[0114]; [0126]-[0127]; [0133]-[0138]; [0169]; [0175]; [0178]-[0181]; [0278]-[0293]; [0220]-[0222]; Fig. 1; Fig. 3; Figs. 5-12)
transmit the distortion-corrected first portion and the distortion-corrected second portion to another electronic device ([0047]; [0054]; [0062]; [0143]; Fig. 2)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 3 is rejected under 35 U.S.C. § 103 as being unpatentable over Kasatani in view of Zhou et. al (US 10911677 B1) (Hereinafter, “Zhou”)
With respect to claim 3, Kasatani teaches the method of claim 2.
Kasatani does not teach the further limitations of claim 2.
However, Zhou, in the same field of endeavor of video streaming, teaches:
wherein the video image stream comprises a stabilized video image stream (Col. 1, lines 36-51, “Many modern electronic image capture devices have multiple cameras that are capable of capturing and producing stabilized video image streams”; Col. 1, lines 5-11 “it relates to techniques for improved stabilization of video image streams…”)
It would have been obvious to one of ordinary skill in the art as of the effective filing date of the claimed invention to modify Kasatani to include the limitations of video stream stabilization, as taught by Zhou. Doing so would predictably improve the efficiency of image operations by ensuring that the base input is visible and stable. The systems readily integrate, as complements.
Claims 8-9 are rejected under 35 U.S.C. § 103 as being unpatentable over Kasatani in view of Chen et. al Pixelwise Deep Sequence Learning for Moving Object Detection (Hereinafter, “Chen”)
With respect to claim 8, Kasatani teaches the method of claim 1.
Kasatani does not explicitly teach the further limitations of claim 8.
However, Chen, in the same field of endeavor of video stream processing, teaches:
wherein the second portion of the first image is detected automatically using a detector ([Abstract]; Fig. 3; Fig. 5)
It would have been obvious to one of ordinary skill in the art as of the effective filing date of the claimed invention to modify Kasatani to include the limitations of detector usage, as taught by Chen. Kasatani would predictably benefit from automation of the clipping process, for advantages of efficiency and potentially improved recognition. The systems readily integrate, as Kasatani otherwise includes a manual masking process for clipping, as described in [0278]-[0293] and Fig. 20.
With respect to claim 9, Kasatani/Chen teaches:
The method of claim 8, wherein the detector comprises a deep neural network (DNN) trained to locate particular surfaces in images (Chen, [Abstract]; Chen, Fig. 3; Chen, Fig. 5)
Claims 11, 16, and 18 are rejected under 35 U.S.C. § 103 as being unpatentable over Kasatani in view of Roulet et. al (US 20200394770 A1) (Hereinafter, “Roulet”)
With respect to claim 11, Kasatani teaches the method of claim 1.
Kasatani does not explicitly teach:
wherein the portion of the surface comprises two or more non-contiguous regions in the first image
However, Roulet, in the same field of endeavor of distortion correction, teaches:
wherein the portion of the surface comprises two or more non-contiguous regions in the first image (Fig. 7; Fig. 8; [0010], “If the same original image also contains buildings, the adaptive dewarping algorithm will apply a different dewarping on the building to keep the lines straight. There is no limit according to the present invention to the types of objects that can be recognized by the adaptive dewarping method and to the dewarping to be applied on them. The dewarping to apply can be defined in advance as a preset for a specific object type, e.g. human face, building, etc. or can be calculated to respect object well-known characteristics such as face proportions, human body proportions, etc. This method is applied on each segmented layer and depth layer, including for the background layer. The background layer consists of objects far from the camera and not having a preset distortion dewarping. In preferred embodiments according to the present invention, the background is dewarped in order to keep the perspective of the scene undistorted. Compared to existing prior art, the adaptive method allows to apply a different dewarping to each kind of object based on type, layer, depth, size, texture, etc”
With respect to the claimed invention, the “non-contiguous regions” are understood to refer to, for example, papers on a desk or wall within the image ([0031] of the specification). The technique of Roulet is predictably capable of segmenting such objects, if they were present atop the desk of Kasatani, Fig. 6. These objects could then receive adaptive distortion accordingly under the pipeline of Roulet, Fig. 7. However, under the broad recitation, simply the introduction of certain objects would be enough to modify Kasatani as teaching the claim (Roulet, [0010], “There is no limit according to the present invention to the types of objects that can be recognized by the adaptive dewarping method and to the dewarping to be applied on them”). The claim does not require any region-specific processing, such that if the requisite papers were simply included explicitly within the image of Kasatani Fig. 6, Kasatani would have anticipated accordingly)
PNG
media_image4.png
1067
777
media_image4.png
Greyscale
It would have been obvious to one of ordinary skill in the art as of the effective filing date of the claimed invention to modify Kasatani to include the limitations of additional objects, as taught by Roulet. Doing so would allow for even more specific dewarping processes, furthering the objective of Kasatani. The systems readily integrate, as it is reasonably expected that the portion of the table included in Kasatani could include further “non-contiguous regions”.
With respect to claim 16, Kasatani teaches the method of claim 14.
Kasatani does not explicitly teach the further limitations of claim 14.
However, Roulet teaches:
wherein the second distortion correction operation is further based on an estimated orientation of the portion of the surface in the first image ([0016]; [0040] “ Also, in some other embodiments according to the present invention, the further processing of the multiple segmented layers before merging them together could also include a perspective tilt correction of at least one element in a segmented layer in order to correct the perspective with respect either to the horizon or any target direction in the scene. This perspective tilt correction is especially useful when the element is a building captured with an unpleasantly looking tilt angle in the original wide-angle image to correct its shape to appears as if it was captured without this tilt angle, but could be applied to any kind of element”; [0046]-[0051] “Euler angles (α,θ) are calculated from the center position Pin0 according to optical distortion using a function F”)
It would have been obvious to one of ordinary skill in the art as of the effective filing date of the claimed invention to modify Kasatani to include the limitations of object orientation, as taught by Roulet. Doing so would allow for more precise and object-specific distortion corrrections. The systems readily integrate, as Kasatani already encourages a multitude of differing distortion operations with respect to chosen areas, and the methods of Roulet would predictably improve this function.
With respect to claim 18, Kasatani teaches the method of claim 1.
Kasatani does not explicitly teach the further limitations of claim 18.
However, Roulet teaches:
performing an occlusion detection operation on the second portion ([0040] “One such example of a relative depth measurement, in no way limiting the possible methods to rank the depth of the layers, is the relative depth measurement based on superposition. In the image 700, the head of the person 703 partially hides the tree 704 because of the superposition, allowing the depth estimation algorithm to rank the relative depth of the tree 704 as being farther away than the person 703 even if absolute distance are not available”; Fig. 8)
excluding at least one detected occlusion from the second distortion correction operation (Fig. 8; “In this example, since the original image had a segmentation layer with people, a custom dewarping with body and face protection based on the context 760 specifically for people will be applied on layers 720 and 725 to get respectively the dewarped layers 765 and 770. This custom dewarping for people does not try to keep the perspective or straight lines of the objects, but rather to keep the shape of human visually pleasing no matter where they are in the field of view. Next, the custom dewarping for unknown or unrecognized objects is applied to layer 730 to get dewarped layer 775. This custom dewarping improve the shape of objects toward the edge of the image as if they were imaged in the center of the image based on the difference of magnification from one edge to the other edge of the object, but without specific correction as for known objects that require a specific correction (building, people). Next, the adaptive dewarping is applied on the layer 740 to get the dewarped view of the building 780. For buildings, it is important for the image to be visually pleasant to keep straight lines and hence the projection applied on this layer keeps the lines straight. Finally, the background layer 750 can also optionally be dewarped if required to get the desired projection, obtaining the dewarped layer 785”; [0041]; [0055] “In order to keep the field of view as close as possible between the original image and the output image, the special dewarping in the corners can ignore the segmentation or the depth layers and not apply a specific dewarping based on context in that zone in the corners of the image. The choice of not applying the specific adaptive dewarping based on context and depth layer in that zone or in another zone of the image or on a specific layer for any other reason is also possible according to the present invention”)
PNG
media_image5.png
1061
788
media_image5.png
Greyscale
It would have been obvious to one of ordinary skill in the art as of the effective filing date of the claimed invention to modify Kasatani to include the limitations of occlusion detection, as taught by Roulet. Doing so would predictably improve the main objective of Kasatani; by increasing the specificity of the distortion correction further. The systems readily integrate, as Kasatani has a similar feature which could predictably be used to meet the limitations of the claim (Kasatani, [0285]-[0286] “More specifically, with reference to FIG. 20, an image 200 of the operator's side is displayed on the (touch panel) liquid crystal display unit 101. When the operator encloses areas such as a circle A1, an ellipse A2, and a pentagon A3, using the tip of his finger and the like and specifies an area M as a masked area excluding the enclosed areas above in the image 200, through a key operation, the CPU identifies the area M as a masked pattern”; Kasatani, Fig. 20)
Additional References
Additionally cited references (see attached PTO-892) otherwise not relied upon above have been made of record in view of the manner in which they evidence the general state of the art.
Inquiry
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NOAH WILLIAM BOYAR whose telephone number is 571-272-8392. The examiner can normally be reached 10:00 – 6:00 EST, Monday – Friday.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chan Park can be reached at 571-272-7409. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NOAH W BOYAR/Examiner, Art Unit 2669
/CHAN S PARK/Supervisory Patent Examiner, Art Unit 2669