DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-20 are pending under this Office action.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 4, 6, 16 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Siver, etc. (US 20180213255 A1) in view of Rondao, etc. (US 20150117518 A1), further in view of Hollis, etc. (US 20030184556 A1).
Regarding claim 1, Siver teaches that an electronic device (See Siver: Fig. 1, and [0072], “Referring generally to FIG. 1, an example system 0100 is shown, having inputs 0110, such as recording device 0111 in communication with a processing device 0130. Recording device 0111 may be a camera, or other image capturing device. In some embodiments, recording device 0111 may be a data capturing device or sensor such as a laser based radar (LIDAR) device. Recording device 0111 may capture image data including colour data 0112 and depth data 0113, also known as volumetric image data. The colour data 0113 may be captured as a series of still images or frames. In some embodiments, recording device 0111 may further capture additional data such as sound data or heat data. In some embodiments, recording device 0111 may comprise multiple recording devices or sensors, such as a first recording device for capturing image data, and a second recording device for recording depth data”) comprising:
memory comprising one or more storage media storing instructions (See Siver: Fig. 1, and [0073], “Data captured by recording device 0111 is passed to processing device 0130. Processing device 0130 may typically be a desktop or laptop computer. In some embodiments, processing device 0130 may be a tablet or handheld computing device. In some embodiments, processing device 0130 may be a computing system, and may include a server or server system, and/or a number of processing devices in communication over a network. Processing device 0130 may contain an input module 0131, memory 0140, a CPU 0132, display 0133, and a network connection 0134”); and
at least one processor comprising processing circuitry(See Siver: Fig. 1, and [0051], “wherein the at least one processor has access to the memory and the memory comprises the computer-readable medium of some other embodiments”),
wherein the instructions, when executed by the at least one processor individually or collectively, cause the electronic (See Siver: Fig. 1, and [0075], “Executing program data may cause display 0133 to show data in the form of a user interface. CPU 0132 may be one or more data processors for executing instructions, and may include one or more of a microcontroller-based platform, a suitable integrated circuit, a GPU or other graphics processing units, and one or more application-specific integrated circuits (ASIC's). CPU 0132 may include an arithmetic logic unit (ALU) for mathematical and/or logical execution of instructions, such operations performed on the data stored in the internal registers. Display 0133 may be a liquid crystal display, a plasma screen, a cathode ray screen device or the like, and may comprise a plurality of separate displays. In some embodiments, display 0133 may include a head mounted display (HMD). Processing device 0130 may have one or more buses (not shown) to facilitate the transfer of data signals between components. The buses may be wire or optical fibre buses in some embodiments, and may be arranged to facilitate parallel and/or bit serial connections”) device to:
obtain depth information of a visual object (See Siver: Fig. 1, and [0072], “Referring generally to FIG. 1, an example system 0100 is shown, having inputs 0110, such as recording device 0111 in communication with a processing device 0130. Recording device 0111 may be a camera, or other image capturing device. In some embodiments, recording device 0111 may be a data capturing device or sensor such as a laser based radar (LIDAR) device. Recording device 0111 may capture image data including colour data 0112 and depth data 0113, also known as volumetric image data. The colour data 0113 may be captured as a series of still images or frames. In some embodiments, recording device 0111 may further capture additional data such as sound data or heat data. In some embodiments, recording device 0111 may comprise multiple recording devices or sensors, such as a first recording device for capturing image data, and a second recording device for recording depth data”; [0120], “Depth is typically recorded as a fixed distance value from a known capture point. Some schemes, such as a reciprocal compression (as opposed to linear, or gamma based) represent closer depth values with greater accuracy. This closely mirrors how most sensors can generate data, losing accuracy over distance. However, sensor accuracy varies widely between capture methods. Laser scanners are capable of capturing depth measurements to the nanometer level, where various market sensors capture between the millimeter and centimeter accuracies, with some surveying style equipment capturing decimeters to meter accuracies”; and [0125], “Another scheme involves having the colour sensor assuming the same position as the depth sensor, just not at the same time. Instead, the colour data is captured from a sensor in a particular location, then the depth sensor is positioned in the same location to capture the depth data. In other words, the colour data and the depth data are spatially correlated, but not temporally correlated. Subsequently, the position and alignment based matching is performed on the two sets of captured data, as is common for laser scanners. However these techniques are inappropriate for temporarily variant subjects, such as moving or flickering objects, and is accordingly better suited and typically used for still capture only”. Note that the object in the captured images of the scene is mapped to the visual object),
identify, based on the depth information of the visual object, depth values (See Siver: Figs. 1-3, and [0072], “Referring generally to FIG. 1, an example system 0100 is shown, having inputs 0110, such as recording device 0111 in communication with a processing device 0130. Recording device 0111 may be a camera, or other image capturing device. In some embodiments, recording device 0111 may be a data capturing device or sensor such as a laser based radar (LIDAR) device. Recording device 0111 may capture image data including colour data 0112 and depth data 0113, also known as volumetric image data. The colour data 0113 may be captured as a series of still images or frames. In some embodiments, recording device 0111 may further capture additional data such as sound data or heat data. In some embodiments, recording device 0111 may comprise multiple recording devices or sensors, such as a first recording device for capturing image data, and a second recording device for recording depth data”; and [0140], “The packing (merger) operation 0330, correspondingly, shows the reverse operation to that shown by unpacking operation 0320. CPU 0132 retrieves input depth 0332 and colour 0331 from original images or image sequences and combines these by performing the packing operation 0330, to prepare the images for input to the container encoder 0340. In the illustrated embodiment, depth information 0332 has a fixed offset 0339 applied by CPU 0132 executing the offset application module 0334 of the processing application 0141. In some embodiments, CPU 0132 also applies packing parameters to the depth information 0334 at step 0336, in order to scale or pack the depth information 0334. In some embodiments, corresponding offset application module 0333 of processing application 0141 may be executed by CPU 0132 to apply a fixed offset to the colour information 0331. In some embodiments, where the colourspace has been converted or otherwise made to accommodate depth information 0332, a colour conversion phase 0335 may be applied by CPU 0132 to translate colour information 0331 to a desired form for storage”),
add the depth values to an alpha channel of an image representing the visual object (See Siver: Figs. 13-14, and [0183], “Referring generally to FIG. 14, a block diagram 4100 is shown illustrating how a depth image 4110 and a colour image 4120 may be packaged into a new container 4140, where the container is configured to support an auxiliary channel. In the illustrated embodiment, container 4140 is able to contain depth 4110 as a virtual, flattened channel. Before depth 4110 is passed into container 4140, it is transformed by transformation process 4130 to prepare it for insertion into the auxiliary or alpha channel of container 4140. Transformation process 4130 may differ depending on the application. For example, transformation process 4130 may include applying a coarse-fine compression, as described above with reference to FIG. 13. In some embodiments, transformation process 4130 may include a transform, rotate, scale (TRS) matrix operation to convert between input and destination coordinates. A further example of transformation process is described below with reference to FIG. 15, where transformation process 4130 may be transformation operation 4200”), and
generate the image (See Siver: Fig. 1, and [0103], “Processing device 0130 may be further configured to communicate with external device 0150 via network connection 0134. In some embodiments, input colour data 0112 and depth data 0113 may be processed by processing application 0141 in order to produce output colour and depth data 0120, which may be sent by processing device 0130 to one or more external devices 0150 via network connection 0134. The processing of colour data 0112 and depth data 0113 may include encoding and packing of the data into an image container. According to some embodiments, the data format of the image container that the data is packed into may be configured for colour data storage. According to some embodiments, the data format of the image container that the data is packed into may be configured for colour data and auxiliary data storage”), and
wherein the alpha channel comprises the depth values (See Siver: Fig. 1, and [0109], “Image formats like .TGA and .PNG can support auxiliary channels, for example by containing specifications to optionally employ a 32 bit arrangement as 4 channels of 8 bits each. This covers R, G, B and A (Alpha). The A is traditionally used for transparency, or exposure. Some image formats may support multiple auxiliary channels. Some image formats may support between 1 and 40 channels or components of colour data and between 1 and 40 channels or components of auxiliary data. In some embodiments, where multiple recording devices are used to capture the data, each recording device may contribute a number of colour and/or depth components to each image. Occasional video methods, such as Theora, and various intermediary formats used in video authoring, will allow for auxiliary channels. In practice, however, support is limited and prone to variation in both performance and hardware compatibility”) and transparencies of the visual object.
However, Siver fails to explicitly disclose that wherein the alpha channel comprises the depth values and transparencies of the visual object.
However, Rondao teaches that wherein the alpha channel comprises the depth values and transparencies of the visual object (See Rondao: Fig. 3, and [0058], “As illustrated in FIG. 3, an associated decoding apparatus 300 in accordance with the present invention, which may for example be provided in a set-top box, a home gateway, a client computer program, or integrated into visualization hardware such as a TV set, comprises a first decoder 350 for decoding portions of the transparency information channel encoded according to a lossless vector graphics encoding scheme (e.g., SVG), a second decoder 380 for decoding portions of the transparency information channel encoded according to a mathematical representation encoding scheme, a third decoder 390 for decoding portions of the transparency information channel encoded according to an MPEG-based encoding scheme, and a detector 310 arranged to detect the encoding applied to specific portions of the transparency information channel, and to submit these portions to the appropriate decoder. The decoded transparency information is provided to the renderer 399”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Siver to have wherein the alpha channel comprises the depth values and transparencies of the visual object as taught by Rondao in order to improve the compression of the alpha component for transparency/opacity signals (See Rondao: Fig. 1, and [0034], “Embodiments of the present invention present an extension of a typical MPEG encoder to provide hybrid encoding of the alpha channel in order to improve the compression of the alpha component for transparency/opacity signals. The main idea is to select the best coding scheme for the alpha channel(s) input(s), which includes optionally representing transparency or depth with non-MPEG compliant codecs, while the YUV channels are encoded using a MPEG-compliant encoding”). Siver teaches a method and system foe image processing that may pack and insert depth into the auxiliary alpha channel; while Rondao teaches a system and method that may treat alpha channel as usable for transparency and depth for hybrid encoding of the alpha channel specifically for transparency and depth information. Therefore, it is obvious for one of ordinary skill in the art to modify Siver by Rondao to have the alpha channel comprise depth values and transparencies. The motivation to modify Siver by Rondao is “Use of known technique to improve similar devices (methods, or products) in the same way”.
However, Siver, modified by Rondao, fails to explicitly disclose that wherein the alpha channel comprises the depth values and transparencies of the visual object.
However, Hollis teaches that wherein the alpha channel comprises the depth values and transparencies of the visual object (See Hollis: Figs. 3A-B, and [0059], “The format shown in FIGS. 3A and 3B may be selected, for example, by specifying a format parameter in a graphics command directed to texture unit 122 for initializing a new texture object. Any given texture mapping will generally have a single overall format--but in this particular example, the two alternate formats shown in FIGS. 3A and 3B are both encompassed by the same format parameter. The most significant bit (bit 15) within the format encoding specifies whether the particular instance of the format contains five bits each of red, green and blue information (RGB5); or alternatively, four bits each of red, green and blue plus three bits of alpha (RGB4A3)”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Siver to have wherein the alpha channel comprises the depth values and transparencies of the visual object as taught by Hollis in order to be of compact coding format (See Hollis: Fig. 1, and [0018], “The present invention takes advantage of this observation by providing, in one particular implementation, a compact image element encoding format that selectively allocates bits on an element-by-element basis to encode multi-bit alpha resolution. This technique may be advantageously used to allocate encoding bits within some image elements for modeling semi-transparency while using those same bits for other purposes (e.g., higher color resolution) in other image elements not requiring a semi-transparency value (e.g., for opaque image elements). Applications include but are not limited to texture mapping in a 3D computer graphics system such as a home video game system or a personal computer”). Siver teaches a method and system for image processing that may pack and insert depth into the auxiliary alpha channel; while Hollis teaches a system and method that may allocate bits within a channel field so that multi-bit transparency information and other data share the same fixed-width field with an indicator or selective encoding so that the alpha channel comprises both the depth values and the transparencies of the visual object. Therefore, it is obvious for one of ordinary skill in the art to modify Siver by Hollis to allocate the bits for the alpha channel bit fields so that the alpha channel comprises depth values and transparencies of the visual object. The motivation to modify Siver by Hollis is “Use of known technique to improve similar devices (methods, or products) in the same way”.
Regarding claim 4, Siver, Rondao and Hollis teach all the features with respect to claim 1 as outlined above. Further, Hollis teaches that the electronic device of claim 1, wherein a first bit sequence indicating the depth values in the alpha channel is positioned after a least significant bit (L SB) of a second bit sequence indicating the transparencies in the alpha channel (See Hollis: Figs. 3A-B, and [0050], “FIGS. 3A and 3B show an example image element variable bit encoding format. In the particular example shown, the format has a fixed length of 16 bits, but how those bits are allocated can vary on an instance-by-instance basis such that the same image map can use different encodings for different elements. In more detail, when the most significant bit (bit 15) is set, the remainder of the format encodes higher resolution color information (for example, five bits each of red, green and blue color values) and defines an opaque image element. When the most significant bit is not set, the format provides lower resolution color information (for example, four bits each of red, green and blue) along with three bits of alpha information defining multiple levels of semi-transparency”. Note that the alpha is placed in the bits 12-14, the most significant bit, and the RGB or depth values is placed in the bit position 0-11, the LSB bit position, this is just the opposite of the claim cited limitation).
Regarding claim 6, Siver, Rondao and Hollis teach all the features with respect to claim 1 as outlined above. Further, Siver teaches that the electronic device of claim 1, wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to:
generate metadata indicating a bit number in the alpha channel, the metadata being reserved for indicating the depth values (See Siver: Fig. 5 and [0163], “When an offset size equal to the size of the original images 1110 and 1120 is used (i.e. the size of each block of data is the image size X by Y), the result is image 1130, with the new image being twice the height of the original images. This places the depth 1110 and the colour 1120 in a vertical arrangement 1130. In the illustrated embodiments, depth image 1110 is written first, so that the depth sub-image 1132 is above colour sub-image 1134. In other embodiments, colour image 1120 may be written first and the sub-images 1132 and 1134 may appear in the opposite order. The order of the sub-images may be an arrangement parameter 0422 stored in the metadata of container 0240, or stored externally to the container. When the resulting image 1130 is being unpacked, a read block size of half of the resulting image 1130 would be used”), and
generate a fiIe comprising the metadata and the image (See Siver: Fig. 5, and [0164], “If the offset is equal to one pixel, the result is image 1150, which results in a horizontal interlacing of the two sets of information. If the offset is set to the image height, vertical interlacing occurs. In the illustrated embodiments, colour image 1110 is written first, so that the colour component appears first. In other embodiments, depth image 1120 may be written first and the colour and depth components may appear in the opposite order. The order of the image components may be an arrangement parameter 0422 stored in the metadata of container 0240, or stored externally to the container”; and [0138], “CPU 0132 transforms input coordinate 0321 using a fixed offset 0327 (which may be stored externally to container 0240, such as in a filename, calibration file, mark-up file, or other associated file, or in the metadata of an associated file), in order to allow data to be read from the correct location. The fixed offset 0327 should be the same offset as that used to encode the original image. In some embodiments, the coordinate + offset 0322 is applied to the depth information 0329 in order to retrieve this information from the original image. In some other embodiments, the offset may instead be applied to the colour information 0328”).
Regarding claim 16, Siver, Rondao and Hollis teach all the features with respect to claim 1 as outlined above. Further, Siver, Rondao and Hollis teach that a method of an electronic device, the method (See Siver: Fig. 1, and [0072], “Referring generally to FIG. 1, an example system 0100 is shown, having inputs 0110, such as recording device 0111 in communication with a processing device 0130. Recording device 0111 may be a camera, or other image capturing device. In some embodiments, recording device 0111 may be a data capturing device or sensor such as a laser based radar (LIDAR) device. Recording device 0111 may capture image data including colour data 0112 and depth data 0113, also known as volumetric image data. The colour data 0113 may be captured as a series of still images or frames. In some embodiments, recording device 0111 may further capture additional data such as sound data or heat data. In some embodiments, recording device 0111 may comprise multiple recording devices or sensors, such as a first recording device for capturing image data, and a second recording device for recording depth data”) comprising:
obtaining depth information of a visual object (See Siver: Fig. 1, and [0072], “Referring generally to FIG. 1, an example system 0100 is shown, having inputs 0110, such as recording device 0111 in communication with a processing device 0130. Recording device 0111 may be a camera, or other image capturing device. In some embodiments, recording device 0111 may be a data capturing device or sensor such as a laser based radar (LIDAR) device. Recording device 0111 may capture image data including colour data 0112 and depth data 0113, also known as volumetric image data. The colour data 0113 may be captured as a series of still images or frames. In some embodiments, recording device 0111 may further capture additional data such as sound data or heat data. In some embodiments, recording device 0111 may comprise multiple recording devices or sensors, such as a first recording device for capturing image data, and a second recording device for recording depth data”; [0120], “Depth is typically recorded as a fixed distance value from a known capture point. Some schemes, such as a reciprocal compression (as opposed to linear, or gamma based) represent closer depth values with greater accuracy. This closely mirrors how most sensors can generate data, losing accuracy over distance. However, sensor accuracy varies widely between capture methods. Laser scanners are capable of capturing depth measurements to the nanometer level, where various market sensors capture between the millimeter and centimeter accuracies, with some surveying style equipment capturing decimeters to meter accuracies”; and [0125], “Another scheme involves having the colour sensor assuming the same position as the depth sensor, just not at the same time. Instead, the colour data is captured from a sensor in a particular location, then the depth sensor is positioned in the same location to capture the depth data. In other words, the colour data and the depth data are spatially correlated, but not temporally correlated. Subsequently, the position and alignment based matching is performed on the two sets of captured data, as is common for laser scanners. However these techniques are inappropriate for temporarily variant subjects, such as moving or flickering objects, and is accordingly better suited and typically used for still capture only”. Note that the object in the captured images of the scene is mapped to the visual object);
identifying, using the depth information, depth values (See Siver: Figs. 1-3, and [0072], “Referring generally to FIG. 1, an example system 0100 is shown, having inputs 0110, such as recording device 0111 in communication with a processing device 0130. Recording device 0111 may be a camera, or other image capturing device. In some embodiments, recording device 0111 may be a data capturing device or sensor such as a laser based radar (LIDAR) device. Recording device 0111 may capture image data including colour data 0112 and depth data 0113, also known as volumetric image data. The colour data 0113 may be captured as a series of still images or frames. In some embodiments, recording device 0111 may further capture additional data such as sound data or heat data. In some embodiments, recording device 0111 may comprise multiple recording devices or sensors, such as a first recording device for capturing image data, and a second recording device for recording depth data”; and [0140], “The packing (merger) operation 0330, correspondingly, shows the reverse operation to that shown by unpacking operation 0320. CPU 0132 retrieves input depth 0332 and colour 0331 from original images or image sequences and combines these by performing the packing operation 0330, to prepare the images for input to the container encoder 0340. In the illustrated embodiment, depth information 0332 has a fixed offset 0339 applied by CPU 0132 executing the offset application module 0334 of the processing application 0141. In some embodiments, CPU 0132 also applies packing parameters to the depth information 0334 at step 0336, in order to scale or pack the depth information 0334. In some embodiments, corresponding offset application module 0333 of processing application 0141 may be executed by CPU 0132 to apply a fixed offset to the colour information 0331. In some embodiments, where the colourspace has been converted or otherwise made to accommodate depth information 0332, a colour conversion phase 0335 may be applied by CPU 0132 to translate colour information 0331 to a desired form for storage”);
adding the depth values to an alpha channel of an image representing the visual object, and generating the image (See Siver: Figs. 13-14, and [0183], “Referring generally to FIG. 14, a block diagram 4100 is shown illustrating how a depth image 4110 and a colour image 4120 may be packaged into a new container 4140, where the container is configured to support an auxiliary channel. In the illustrated embodiment, container 4140 is able to contain depth 4110 as a virtual, flattened channel. Before depth 4110 is passed into container 4140, it is transformed by transformation process 4130 to prepare it for insertion into the auxiliary or alpha channel of container 4140. Transformation process 4130 may differ depending on the application. For example, transformation process 4130 may include applying a coarse-fine compression, as described above with reference to FIG. 13. In some embodiments, transformation process 4130 may include a transform, rotate, scale (TRS) matrix operation to convert between input and destination coordinates. A further example of transformation process is described below with reference to FIG. 15, where transformation process 4130 may be transformation operation 4200”),
wherein the alpha channel (See Siver: Fig. 1, and [0109], “Image formats like .TGA and .PNG can support auxiliary channels, for example by containing specifications to optionally employ a 32 bit arrangement as 4 channels of 8 bits each. This covers R, G, B and A (Alpha). The A is traditionally used for transparency, or exposure. Some image formats may support multiple auxiliary channels. Some image formats may support between 1 and 40 channels or components of colour data and between 1 and 40 channels or components of auxiliary data. In some embodiments, where multiple recording devices are used to capture the data, each recording device may contribute a number of colour and/or depth components to each image. Occasional video methods, such as Theora, and various intermediary formats used in video authoring, will allow for auxiliary channels. In practice, however, support is limited and prone to variation in both performance and hardware compatibility”) comprises the depth values (See Rondao: Fig. 3, and [0058], “As illustrated in FIG. 3, an associated decoding apparatus 300 in accordance with the present invention, which may for example be provided in a set-top box, a home gateway, a client computer program, or integrated into visualization hardware such as a TV set, comprises a first decoder 350 for decoding portions of the transparency information channel encoded according to a lossless vector graphics encoding scheme (e.g., SVG), a second decoder 380 for decoding portions of the transparency information channel encoded according to a mathematical representation encoding scheme, a third decoder 390 for decoding portions of the transparency information channel encoded according to an MPEG-based encoding scheme, and a detector 310 arranged to detect the encoding applied to specific portions of the transparency information channel, and to submit these portions to the appropriate decoder. The decoded transparency information is provided to the renderer 399”) and transparencies of the visual object (See Hollis: Figs. 3A-B, and [0059], “The format shown in FIGS. 3A and 3B may be selected, for example, by specifying a format parameter in a graphics command directed to texture unit 122 for initializing a new texture object. Any given texture mapping will generally have a single overall format--but in this particular example, the two alternate formats shown in FIGS. 3A and 3B are both encompassed by the same format parameter. The most significant bit (bit 15) within the format encoding specifies whether the particular instance of the format contains five bits each of red, green and blue information (RGB5); or alternatively, four bits each of red, green and blue plus three bits of alpha (RGB4A3)”).
Regarding claim 19, Siver, Rondao and Hollis teach all the features with respect to claim 16 as outlined above. Further, Hollis teaches that the method of claim 16, wherein a first bit sequence indicating the depth values in the alpha channel is positioned after a least significant bit (LSB) of a second bit sequence indicating the transparencies in the alpha channel (See Hollis: Figs. 3A-B, and [0050], “FIGS. 3A and 3B show an example image element variable bit encoding format. In the particular example shown, the format has a fixed length of 16 bits, but how those bits are allocated can vary on an instance-by-instance basis such that the same image map can use different encodings for different elements. In more detail, when the most significant bit (bit 15) is set, the remainder of the format encodes higher resolution color information (for example, five bits each of red, green and blue color values) and defines an opaque image element. When the most significant bit is not set, the format provides lower resolution color information (for example, four bits each of red, green and blue) along with three bits of alpha information defining multiple levels of semi-transparency”. Note that the alpha is placed in the bits 12-14, the most significant bit, and the RGB or depth values is placed in the bit position 0-11, the LSB bit position, this is just the opposite of the claim cited limitation).
Claims 2-3 and 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Siver, etc. (US 20180213255 A1) in view of Rondao, etc. (US 20150117518 A1), further in view of Hollis, etc. (US 20030184556 A1) and Fleureau, etc. (US 20210385454 A1).
Regarding claim 2, Siver, Rondao and Hollis teach all the features with respect to claim 1 as outlined above. However, Siver, modified by Rondao and Hollis, fails to explicitly disclose that the electronic device of claim 1, wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to: identify a range of the depth values, determine, based on the range of the depth values, a first bit number representing the depth values in the alpha channel, and a second bit number representing the transparencies, and generate, based on the determined first bit number and the determined second bit number, the image.
However, Fleureau teaches that the electronic device of claim 1, wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to:
identify a range of the depth values (See Fleureau: Fig. 5, and [0109], “FIG. 5 shows an example of a picture 50 comprising the depth information of the points of the 3D scene 10, for example in the case wherein the 3D scene is acquired from a single point of view, according to a non-limiting embodiment of the present principles. The picture 50 corresponds to an array of first pixels, each first pixel comprising data representative of depth. The picture 50 may also be called a depth map. The data corresponds for example to a floating-point value indicating, for each first pixel, the radial distance to z to the viewpoint of the picture 50 (or of the center of projection when the picture 50 is obtained by projection). The depth data may be obtained with one or more depth sensors or may be known a priori, for example for CGI parts of the scene. The depth range [z.sub.min, z.sub.max] contained in picture 50, i.e. the range of depth values comprised between the minimal depth value z.sub.min and the maximal depth value z.sub.max of the scene, may be large, for example from 0 to 100 meters. The depth data is represented with a shade of grey in FIG. 5, the darker the pixel (or point), the closer to the viewpoint”),
determine, based on the range of the depth values, a first bit number representing the depth values in the alpha channel, and a second bit number representing the transparencies (See Fleureau: Fi. 1, and [0085], “To reach that aim, data representative of the depth (e.g. distance or depth values expressed as floating-point values associated with the elements, e.g. points, of the 3D scene) of the 3D scene is quantized in a number of quantized depth values that is greater than the number of encoding values allowed by a determined encoding bit depth. For example, 8 bits encoding bit depth allows an encoding with 256 (2.sup.8) values and 10 bits encoding bit depth allows an encoding with 1024 (2.sup.10) values while the number of quantized depth values may for example be equal to 16384 (2.sup.14) or to 65536 (2.sup.16). Quantizing depth data with a great number of values enable the quantizing of big range of depth data (for example depth or distances comprised between 0 and 50 meters or between 0 and 100 meters or even bigger ranges) with a quantization step that remains small over the whole range, minimizing the quantization error, especially for objects close to the point(s) of view of the scene (e.g. foregrounds objects)”; [0086], “The picture comprising the data representative of depth is divided into blocks of pixels (e.g. blocks of 8×8 or 16×16 pixels) and a first set of candidate quantization parameters is determined for each block of pixels. A candidate quantization parameter corresponds to a quantization value that is representative of a range of quantization values associated with the pixels of the block (the quantization values being obtained by quantizing the depth data stored in the pixels of the picture). A first set of candidate quantization parameters is determined considering the number of encoding values for the considered block (e.g. 1024 values for the block), a candidate quantization parameter corresponding for example to a reference quantization value that may be used as a starting value for representing a range of quantized depth value in the limit of the number of encoding values”; and [0087], “A second set of quantization parameters is determined as a subset of the union of the first sets, i.e. the second set comprises a part of the candidate quantization parameters determined for all blocks of pixels of the picture. The second set being determined by retrieving the minimal number of candidate quantization parameters that enables to represent all ranges of quantized depth values of all blocks of pixels of the whole picture, i.e. by retrieving the candidate quantization parameters that are common for several blocks of pixels, when they exist”), and
generate, based on the determined first bit number and the determined second bit number, the image (See Fleureau: Fig. 6, and [0240], “Implementations of the various processes and features described herein may be embodied in a variety of different equipment or applications, particularly, for example, equipment or applications associated with data encoding, data decoding, view generation, texture processing, and other processing of images and related texture information and/or depth information. Examples of such equipment include an encoder, a decoder, a post-processor processing output from a decoder, a pre-processor providing input to an encoder, a video coder, a video decoder, a video codec, a web server, a set-top box, a laptop, a personal computer, a cell phone, a PDA, and other communication devices. As should be clear, the equipment may be mobile and even installed in a mobile vehicle”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Siver to have the electronic device of claim 1, wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to: identify a range of the depth values, determine, based on the range of the depth values, a first bit number representing the depth values in the alpha channel, and a second bit number representing the transparencies, and generate, based on the determined first bit number and the determined second bit number, the image as taught by Fleureau in order to minimize the quantization error for objects close to the point of view of the scene (See Fleureau: Fig. 1, and [0085], “Quantizing depth data with a great number of values enable the quantizing of big range of depth data (for example depth or distances comprised between 0 and 50 meters or between 0 and 100 meters or even bigger ranges) with a quantization step that remains small over the whole range, minimizing the quantization error, especially for objects close to the point(s) of view of the scene (e.g. foregrounds objects)”). Siver teaches a method and system for image processing that may pack and insert depth into the auxiliary alpha channel; while Fleureau teaches a system and method that may use different bit number to quantize the depth values based on the depth values. Therefore, it is obvious for one of ordinary skill in the art to modify Siver by Fleureau to quantize the depth values using different number of bit to minimize the quantization error. The motivation to modify Siver by Fleureau is “Use of known technique to improve similar devices (methods, or products) in the same way”.
Regarding claim 3, Siver, Rondao, Hollis and Fleureau teach all the features with respect to claim 2 as outlined above. Further, Hollis teaches that the electronic device of claim 2, wherein, in a first case that the depth values are represented by first bits of the first bit number, the first bits being determined using the depth information, the transparencies are represented by second bits of the second bit number, which are subtracted from a total number of bits in the alpha channel by the first bits of the first bit number (See Hollis: Figs. 3A-B, and [0050], “FIGS. 3A and 3B show an example image element variable bit encoding format. In the particular example shown, the format has a fixed length of 16 bits, but how those bits are allocated can vary on an instance-by-instance basis such that the same image map can use different encodings for different elements. In more detail, when the most significant bit (bit 15) is set, the remainder of the format encodes higher resolution color information (for example, five bits each of red, green and blue color values) and defines an opaque image element. When the most significant bit is not set, the format provides lower resolution color information (for example, four bits each of red, green and blue) along with three bits of alpha information defining multiple levels of semi-transparency”), and
wherein, in a second case that the depth values are represented by third bits of a third bit number that is greater than the first bit number, the third bits being determined using the depth information, the transparencies are represented by fourth bits of a fourth bit number, which are subtracted from the total number of bits in the alpha channel by the third bits of the third bit number (See Hollis: Figs. 3A-B, and [0059], “The format shown in FIGS. 3A and 3B may be selected, for example, by specifying a format parameter in a graphics command directed to texture unit 122 for initializing a new texture object. Any given texture mapping will generally have a single overall format--but in this particular example, the two alternate formats shown in FIGS. 3A and 3B are both encompassed by the same format parameter. The most significant bit (bit 15) within the format encoding specifies whether the particular instance of the format contains five bits each of red, green and blue information (RGB5); or alternatively, four bits each of red, green and blue plus three bits of alpha (RGB4A3)”; and [0012], “While the approach of selecting between single-word RGB format and double-word RGBA format is very useful, it also has certain significant limitations. For example, in resource-constrained 3-D graphics systems such as 3-D home video games, it may be especially important as a practical matter to conserve memory usage and associated memory access time. This might mean, for example, that in the context of a real time interactive game, the programmer may rarely (if ever) have the luxury of activating the double-word RGBA mode because of memory space or speed performance considerations. In other words, even when using a system that provides an alpha mode, the game programmer may sometimes be unable to take advantage of it without degrading image complexity (e.g., number of textures) and/or speed performance”).
Regarding claim 17, Siver, Rondao and Hollis teach all the features with respect to claim 16 as outlined above. Further, Fleureau teaches that the method of claim 16, wherein the identifying, based on the depth information of the visual object, depth values, comprises:
identifying a range of the depth values (See Fleureau: Fig. 5, and [0109], “FIG. 5 shows an example of a picture 50 comprising the depth information of the points of the 3D scene 10, for example in the case wherein the 3D scene is acquired from a single point of view, according to a non-limiting embodiment of the present principles. The picture 50 corresponds to an array of first pixels, each first pixel comprising data representative of depth. The picture 50 may also be called a depth map. The data corresponds for example to a floating-point value indicating, for each first pixel, the radial distance to z to the viewpoint of the picture 50 (or of the center of projection when the picture 50 is obtained by projection). The depth data may be obtained with one or more depth sensors or may be known a priori, for example for CGI parts of the scene. The depth range [z.sub.min, z.sub.max] contained in picture 50, i.e. the range of depth values comprised between the minimal depth value z.sub.min and the maximal depth value z.sub.max of the scene, may be large, for example from 0 to 100 meters. The depth data is represented with a shade of grey in FIG. 5, the darker the pixel (or point), the closer to the viewpoint”),
determining, based on the range of the depth values, a first bit number to represent the depth values in the alpha channel, and a second bit number to represent the transparencies in the alpha channel (See Fleureau: Fi. 1, and [0085], “To reach that aim, data representative of the depth (e.g. distance or depth values expressed as floating-point values associated with the elements, e.g. points, of the 3D scene) of the 3D scene is quantized in a number of quantized depth values that is greater than the number of encoding values allowed by a determined encoding bit depth. For example, 8 bits encoding bit depth allows an encoding with 256 (2.sup.8) values and 10 bits encoding bit depth allows an encoding with 1024 (2.sup.10) values while the number of quantized depth values may for example be equal to 16384 (2.sup.14) or to 65536 (2.sup.16). Quantizing depth data with a great number of values enable the quantizing of big range of depth data (for example depth or distances comprised between 0 and 50 meters or between 0 and 100 meters or even bigger ranges) with a quantization step that remains small over the whole range, minimizing the quantization error, especially for objects close to the point(s) of view of the scene (e.g. foregrounds objects)”; [0086], “The picture comprising the data representative of depth is divided into blocks of pixels (e.g. blocks of 8×8 or 16×16 pixels) and a first set of candidate quantization parameters is determined for each block of pixels. A candidate quantization parameter corresponds to a quantization value that is representative of a range of quantization values associated with the pixels of the block (the quantization values being obtained by quantizing the depth data stored in the pixels of the picture). A first set of candidate quantization parameters is determined considering the number of encoding values for the considered block (e.g. 1024 values for the block), a candidate quantization parameter corresponding for example to a reference quantization value that may be used as a starting value for representing a range of quantized depth value in the limit of the number of encoding values”; and [0087], “A second set of quantization parameters is determined as a subset of the union of the first sets, i.e. the second set comprises a part of the candidate quantization parameters determined for all blocks of pixels of the picture. The second set being determined by retrieving the minimal number of candidate quantization parameters that enables to represent all ranges of quantized depth values of all blocks of pixels of the whole picture, i.e. by retrieving the candidate quantization parameters that are common for several blocks of pixels, when they exist”), and
wherein the generating the image comprises generating the image based on the determined first number and the determined second number . (See Fleureau: Fig. 6, and [0240], “Implementations of the various processes and features described herein may be embodied in a variety of different equipment or applications, particularly, for example, equipment or applications associated with data encoding, data decoding, view generation, texture processing, and other processing of images and related texture information and/or depth information. Examples of such equipment include an encoder, a decoder, a post-processor processing output from a decoder, a pre-processor providing input to an encoder, a video coder, a video decoder, a video codec, a web server, a set-top box, a laptop, a personal computer, a cell phone, a PDA, and other communication devices. As should be clear, the equipment may be mobile and even installed in a mobile vehicle”).
Regarding claim 18, Siver, Rondao, Hollis and Fleureau teach all the features with respect to claim 17 as outlined above. Further, Hollis teaches that the method of claim 17, wherein, in a first case that the depth values are represented by first bits of the first bit number, the first bits being determined using the depth information, the transparencies are represented by second bits of the second bit number, which are subtracted from a total number of bits in the alpha channel by the first bits of the first bit number (See Hollis: Figs. 3A-B, and [0050], “FIGS. 3A and 3B show an example image element variable bit encoding format. In the particular example shown, the format has a fixed length of 16 bits, but how those bits are allocated can vary on an instance-by-instance basis such that the same image map can use different encodings for different elements. In more detail, when the most significant bit (bit 15) is set, the remainder of the format encodes higher resolution color information (for example, five bits each of red, green and blue color values) and defines an opaque image element. When the most significant bit is not set, the format provides lower resolution color information (for example, four bits each of red, green and blue) along with three bits of alpha information defining multiple levels of semi-transparency”), and
wherein, in a second case that the depth values are represented by third bits of a third bit number greater than the first bit number, the third bits being determined using the depth information, the transparencies are represented by fourth bits of a fourth bit number, which are subtracted from the total number of bits in the alpha channel by the third bits of the third bit number (See Hollis: Figs. 3A-B, and [0059], “The format shown in FIGS. 3A and 3B may be selected, for example, by specifying a format parameter in a graphics command directed to texture unit 122 for initializing a new texture object. Any given texture mapping will generally have a single overall format--but in this particular example, the two alternate formats shown in FIGS. 3A and 3B are both encompassed by the same format parameter. The most significant bit (bit 15) within the format encoding specifies whether the particular instance of the format contains five bits each of red, green and blue information (RGB5); or alternatively, four bits each of red, green and blue plus three bits of alpha (RGB4A3)”; and [0012], “While the approach of selecting between single-word RGB format and double-word RGBA format is very useful, it also has certain significant limitations. For example, in resource-constrained 3-D graphics systems such as 3-D home video games, it may be especially important as a practical matter to conserve memory usage and associated memory access time. This might mean, for example, that in the context of a real time interactive game, the programmer may rarely (if ever) have the luxury of activating the double-word RGBA mode because of memory space or speed performance considerations. In other words, even when using a system that provides an alpha mode, the game programmer may sometimes be unable to take advantage of it without degrading image complexity (e.g., number of textures) and/or speed performance”).
Claims 5 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Siver, etc. (US 20180213255 A1) in view of Rondao, etc. (US 20150117518 A1), further in view of Hollis, etc. (US 20030184556 A1), Fleureau, etc. (US 20210385454 A1) and Boyce, etc. (US 20190373241 A1).
Regarding claim 5, Siver, Rondao and Hollis and Fleureau, teach all the features with respect to claim 1 as outlined above. However, Siver, modified by Rondao, Hollis and Fleureau, fails to explicitly disclose that the electronic device of claim 1, wherein a first bit sequence indicating the depth values in the alpha channel is positioned before a most significant bit (M SB) of a second bit sequence indicating the transparencies in the alpha channel.
However. Boyce teaches that the electronic device of claim 1, wherein a first bit sequence indicating the depth values in the alpha channel is positioned before a most significant bit (M SB) of a second bit sequence indicating the transparencies in the alpha channel (See Boyce: Figs. 22A-B, and [0191], “FIG. 22B illustrates one embodiment of bit depth coding logic 2010 being implemented at a video client, at which bit depth coding logic 2010 includes decoder 2041 and reconstruction logic 2220. Decoder 2020 decodes received encoded data into YUV video component data, which is transmitted to reconstruction logic 2220. Reconstruction logic 2220 converts the YUV data back into the original 16 bit depth data. In one embodiment, reconstruction logic 2220 performs the conversion by converting the UV value into the 8 LSB of the bit depth data converted, and converting the Y component into the 8 MSB.”. Note that the bit field is divided into two part, the MSB is assigned to the Y component (luma/monochrome) and LSB is assigned to UV components (chroma), combined with the secondary and the tertiary arts, this explicitly teaching of splitting the bit field into MSB and LSB portion can be mapped to the depth values being represented by MSB, and the transparencies being represented in LSB).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Siver to have the electronic device of claim 1, wherein a first bit sequence indicating the depth values in the alpha channel is positioned before a most significant bit (M SB) of a second bit sequence indicating the transparencies in the alpha channel as taught by Boyce in order to enable an efficient execution environment in a face of higher latency memory accesses (See Boyce: Figs. 6A-B, and [0076], “Each of the execution units 608A-608N is capable of multi-issue single instruction multiple data (SIMD) execution and multi-threaded operation enables an efficient execution environment in the face of higher latency memory accesses”). Siver teaches a method and system for image processing that may pack and insert depth into the auxiliary alpha channel; while Boyce teaches a system and method that may split the alpha channel bit field into two portions, MSB is designed for one data, and LSB is designed for another data. Therefore, it is obvious for one of ordinary skill in the art to modify Siver by Boyce to divide the alpha channel bit field into MSB and LSB, and use MSB to represent depth values and use the LSB to represent transparencies. The motivation to modify Siver by Boyce is “Use of known technique to improve similar devices (methods, or products) in the same way”.
Regarding claim 20, Siver, Rondao and Hollis teach all the features with respect to claim 16 as outlined above. Further, Boyce teaches that the method of claim 16, wherein a first bit sequence indicating the depth values in the alpha channel is positioned before a most significant bit (MSB) of a second bit sequence indicating the transparencies in the alpha channel (See Boyce: Figs. 22A-B, and [0191], “FIG. 22B illustrates one embodiment of bit depth coding logic 2010 being implemented at a video client, at which bit depth coding logic 2010 includes decoder 2041 and reconstruction logic 2220. Decoder 2020 decodes received encoded data into YUV video component data, which is transmitted to reconstruction logic 2220. Reconstruction logic 2220 converts the YUV data back into the original 16 bit depth data. In one embodiment, reconstruction logic 2220 performs the conversion by converting the UV value into the 8 LSB of the bit depth data converted, and converting the Y component into the 8 MSB.”. Note that the bit field is divided into two part, the MSB is assigned to the Y component (luma/monochrome) and LSB is assigned to UV components (chroma), combined with the secondary and the tertiary arts, this explicitly teaching of splitting the bit field into MSB and LSB portion can be mapped to the depth values being represented by MSB, and the transparencies being represented in LSB).
Claims 7 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Siver, etc. (US 20180213255 A1) in view of Rondao, etc. (US 20150117518 A1), further in view of Hollis, etc. (US 20030184556 A1), Fleureau, etc. (US 20210385454 A1), Boyce, etc. (US 20190373241 A1) and Diggins (US 20160217544 A1).
Regarding claim 7, Siver, Rondao and Hollis teach all the features with respect to claim 1 as outlined above. However, Siver, modified by Rondao, Hollis, Fleureau and Boyce, fails to explicitly disclose that the electronic device of claim 1, wherein the image comprises a first area corresponding to the visual object and a second area surrounding the first area, wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to: insert, in first pixels of the alpha channel, the depth values and the transparencies, and insert, in second pixels of the alpha channel, bit numbers of the depth values. inserted in the first pixels, and wherein the first pixels correspond to the first area and the second pixels correspond to the second area.
However, Diggins teaches that the electronic device of claim 1, wherein the image comprises a first area corresponding to the visual object and a second area surrounding the first area (See Diggins: Fig. 4, and [0043], “As explained above, the image is divided into tiles, and all the qualifying pixels in a tile carry the same metadata bit (or bits in other examples). A suitable arrangement of tiles is shown in FIG. 4. The image (40) is divided into 1024 equally-sized tiles, excluding the image edge regions (41), (42), (43) and (44). The horizontal and vertical tile dimensions are chosen to be a convenient multiple of the respective pixel pitches, so that every tile contains the same number of pixels. The tiles are grouped into four equally-sized quadrants (45), (46), (47) and (48). Each quadrant carries the same data so as to increase the likelihood that suitable pixels are available for coding, and to provide redundancy”),
wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device (See Diggins: Figs. 3-6, and [0014], “In a further aspect the invention consists in a method and apparatus for processing a video data stream to decode at least four sets of metadata that respectively describe at least four video fields or frames preceding a current field or frame in which decoded metadata describing a previous field or frame is decoded from a later field or frame and associated with a stored or delayed copy of the said preceding field or frame”) to:
insert, in first pixels of the alpha channel, the depth values and the transparencies (See Diggins: Fig. 1-3, and [0004], “The association between the metadata and its respective image is frequently achieved by carrying the metadata in the blanking intervals of television signals, but, because these intervals are discarded by some video processing equipment, the metadata can be lost, or delayed differently from the video itself. This difficulty can be avoided by encoding the metadata within the active area of the video frames; this technique is known as ‘watermarking’. Typical watermarks are intended to be imperceptible to a viewer; known watermarking algorithms include methods that modify the frequency spectrum of the video, and methods that add low amplitude data signals to the video signal”; and [0029], “Video fields (1) and associated metadata packets (2), comprising metadata applicable to or describing the respective video field, are input to a watermark encoder (3). The metadata packets (2) may be carried in the blanking intervals of the video fields (1), or the association between respective fields and metadata packets may be achieved by some other means, for example a defined time relationship. The metadata packets (2) are also input to an N-stage delay (4) that provides a set of N delayed metadata outputs (5) which correspond to the N metadata packets respectively associated with the N previous fields. As each new video field (1) is input to the system, a corresponding current-field metadata packet (2), and set of N delayed metadata packets (5) is input to the watermark encoder (3)”), and
insert, in second pixels of the alpha channel, bit numbers of the depth values. inserted in the first pixels (See Diggins: Fig. 1-5, and [0010], “Suitably, identical modifications are applied to (preferably the low-significance bits defining the values of) all the pixels of a contiguous set of pixels comprising a spatial region within a field or frame so as to encode one or more bits of metadata within that spatial region of that field or frame”; and [0029], “Video fields (1) and associated metadata packets (2), comprising metadata applicable to or describing the respective video field, are input to a watermark encoder (3). The metadata packets (2) may be carried in the blanking intervals of the video fields (1), or the association between respective fields and metadata packets may be achieved by some other means, for example a defined time relationship. The metadata packets (2) are also input to an N-stage delay (4) that provides a set of N delayed metadata outputs (5) which correspond to the N metadata packets respectively associated with the N previous fields. As each new video field (1) is input to the system, a corresponding current-field metadata packet (2), and set of N delayed metadata packets (5) is input to the watermark encoder (3)”; and [0044], “FIG. 5 shows how the metadata bits are allocated to the tiles. It shows an expanded view of the top left quadrant (45) of FIG. 4, and a few adjacent tiles in the top right and bottom left quadrants. Of the 256 tiles in each quadrant, 192 tiles carry data bits d.sub.1 to d.sub.192 respectively. In the top left quadrant (50), the leftmost tile in the second row (51) carries data bit d.sub.1, and the rightmost tile in the bottom row (52) carries data bit d.sub.192. The remaining 64 tiles are divided into two groups of 32 tiles; one group always carries the value one, and the other group always carries the value zero. These 64 tiles are arranged in a regular quincunx pattern of ones and zeros; see, for example the ‘one’ tile (53) and the ‘zero’ tile (54). These tiles with constant encoded data values are used as ‘reference’ tiles to aid the decoding of the data as will be explained below”), and
wherein the first pixels correspond to the first area and the second pixels correspond to the second area (See Diggins: Fig. 4, and [0043], “As explained above, the image is divided into tiles, and all the qualifying pixels in a tile carry the same metadata bit (or bits in other examples). A suitable arrangement of tiles is shown in FIG. 4. The image (40) is divided into 1024 equally-sized tiles, excluding the image edge regions (41), (42), (43) and (44). The horizontal and vertical tile dimensions are chosen to be a convenient multiple of the respective pixel pitches, so that every tile contains the same number of pixels. The tiles are grouped into four equally-sized quadrants (45), (46), (47) and (48). Each quadrant carries the same data so as to increase the likelihood that suitable pixels are available for coding, and to provide redundancy”. Note that the different spatial regions of the image including central versus edge/surrounding regions) can be treated differently, and different regions are allocated different data values, including metadata bit, Placing this metadata into alpha channel of the second (surrounding) pixels is a direct application o the region-based metadata embedding techniques to the alpha channel including the depth values and the transparencies as disclosed in the primary and the tertiary art).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Siver to have the electronic device of claim 1, wherein the image comprises a first area corresponding to the visual object and a second area surrounding the first area, wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to: insert, in first pixels of the alpha channel, the depth values and the transparencies, and insert, in second pixels of the alpha channel, bit numbers of the depth values. inserted in the first pixels, and wherein the first pixels correspond to the first area and the second pixels correspond to the second area as taught by Diggins in order to prevents the metadata from being lost or damaged by video processing (See Diggins: Fig. 1 and [0006], “There is thus a need to carry metadata associated with fields or frames of a motion-image stream in a robust manner that prevents the metadata from being lost or damaged by video processing”). Siver teaches a method and system for image processing that may pack and insert depth into the auxiliary alpha channel; while Diggins teaches a system and method that may divide the image into different spatial regions, and embed the metadata into different regions to avoid losing the metadata during the image processing. Therefore, it is obvious for one of ordinary skill in the art to modify Siver by Diggins to divide the image into different regions, and embed the metadata and other values including the depth value and transparencies into the different spatial regions of the image. The motivation to modify Siver by Diggins is “Use of known technique to improve similar devices (methods, or products) in the same way”.
Regarding claim 14, Siver, Rondao and Hollis teach all the features with respect to claim 1 as outlined above. Further, Fleureau teaches that the electronic device of claim 1, wherein the visual object comprises an avatar representing a user of the electronic device (Saee Fleureau: Fig. 2, and [0094], “FIG. 2 shows a three-dimension (3D) model of an object 20 and points of a point cloud 21 corresponding to the 3D model 20. The 3D model 20 and the point cloud 21 may for example correspond to a possible 3D representation of an object of the 3D scene 10, for example the head of a character. The model 20 may be a 3D mesh representation and points of point cloud 21 may be the vertices of the mesh. Points of the point cloud 21 may also be points spread on the surface of faces of the mesh. The model 20 may also be represented as a splatted version of the point cloud 21, the surface of the model 20 being created by splatting the points of the point cloud 21. The model 20 may be represented by a lot of different representations such as voxels or splines. FIG. 2 illustrates the fact that a point cloud may be defined with a surface representation of a 3D object and that a surface representation of a 3D object may be generated from a point of cloud. As used herein, projecting points of a 3D object (by extension points of a 3D scene) onto an image is equivalent to projecting any image representation of this 3D object to create an object”. Note that the 3D model of the person is mapped to the avatar of the user).
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Siver, etc. (US 20180213255 A1) in view of Rondao, etc. (US 20150117518 A1), further in view of Hollis, etc. (US 20030184556 A1), Fleureau, etc. (US 20210385454 A1), Boyce, etc. (US 20190373241 A1), Diggins (US 20160217544 A1) and Molyneaux, etc. (US 20130342527 A1).
Regarding claim 15, Siver, Rondao and Hollis teach all the features with respect to claim 14 as outlined above. Further, However, Siver, modified by Rondao, Hollis, Fleureau and Boyce, fails to explicitly disclose that the electronic device of claim 14, wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to, based on receiving an input to render the avatar, start to generate the image using a virtual space comprising the avatar.
However, Molyneaux teaches that the electronic device of claim 14, wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to, based on receiving an input to render the avatar, start to generate the image using a virtual space comprising the avatar (See Molyneaux: Figs. 3-4, and [0034], “The characteristic metrics are also measured in 3D meshes. A linear map is built upon the metrics and shape transform matrices S that are represented by a 9.times.N vector, here N is the number of triangle faces of the template mesh. A principal component analysis (PCA) is performed on this linear space to capture the dominating subspace. Thus, given a set of characteristic metrics--for instance a set including a height, an arm length, a shoulder width, a chest radius and a waist radius--the body shape transform S can be approximated by PCA”; and [0038], “At 50 of method 30, a virtual head mesh distinct from the virtual body mesh is constructed. The virtual head mesh may be constructed based on a second depth map different from the first depth map referred to hereinabove. The second depth map may be acquired when the subject is closer to the depth camera than when the first depth map is acquired. In one embodiment, the second depth map may be a composite of three different image captures of the subject's head: a front view, a view turned thirty degrees to the right, and a view turned thirty degrees to the left. In the second depth map, the subject's facial features may be resolved more finely than in the first depth map. In other embodiments, the subject's head may be rotated by angles greater than or less than thirty degrees between successive image captures”).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Siver to have the electronic device of claim 14, wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to, based on receiving an input to render the avatar, start to generate the image using a virtual space comprising the avatar as taught by Molyneaux in order to construct the high-fidelity, geometrically realistic and seamless head/body mesh avatar in the efficient and the cost effective manner (See Molyneaux: Fig. 2 and [0040], “At 52 the virtual body mesh is connected to the virtual head mesh. In this step, the head of the virtual body template mesh is first cut out, and then is connected to the virtual head template mesh by triangulating the two open boundaries of the template meshes. The connected model is then stored in the system and loaded when the virtual body mesh and the virtual head mesh are ready. The two template meshes are replaced by two virtual meshes, respectively, since they have the same connectivities. The scale of the virtual head mesh is adjusted according to the proportion consist ent with the virtual body mesh. The vertices around the neck are also smoothed, while the other vertices are held fixed. In this manner, a geometrically realistic and seamless head/body mesh may be constructed”). Siver teaches a method and system for image processing that may pack and insert depth into the auxiliary alpha channel; while Molyneaux teaches a system and method that may divide the image into different spatial regions, and embed the metadata into different regions to avoid losing the metadata during the image processing. Therefore, it is obvious for one of ordinary skill in the art to modify Siver by Molyneaux to divide the image into different regions, and embed the metadata and other values including the depth value and transparencies into the different spatial regions of the image. The motivation to modify Siver by Molyneaux is “Use of known technique to improve similar devices (methods, or products) in the same way”.
Allowable Subject Matter
Claims 8-9 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The best arts searched, Siver, etc. (US 20180213255 A1), Rondao, etc. (US 20150117518 A1), Hollis, etc. (US 20030184556 A1), Fleureau, etc. (US 20210385454 A1), Boyce, etc. (US 20190373241 A1) and Diggins (US 20160217544 A1), do not teach the cited limitations of “ the electronic device of claim 1, wherein the image is a first image, and wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to: obtain other depth information based on a shape of the visual object at a second moment after a first moment corresponding to the first image, generate a second image having the alpha channel comprising differences between other depth values indicated by the other depth information and the depth values in the alpha channel of the first image, and generate a video comprising the first image and the second image.”
Claim 10 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The best arts searched, Siver, etc. (US 20180213255 A1), Rondao, etc. (US 20150117518 A1), Hollis, etc. (US 20030184556 A1), Fleureau, etc. (US 20210385454 A1), Boyce, etc. (US 20190373241 A1) and Diggins (US 20160217544 A1), do not teach the cited limitations of “the electronic device of claim 1, wherein the image is a first image, and wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to, based on a preset number of images which are rendered after the first image corresponding to a key frame within a video, generate the preset number of images comprising the alpha channel comprising difference values with respect to the depth values in the alpha channel of the first image.”
Claim 11 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The best arts searched, Siver, etc. (US 20180213255 A1), Rondao, etc. (US 20150117518 A1), Hollis, etc. (US 20030184556 A1), Fleureau, etc. (US 20210385454 A1), Boyce, etc. (US 20190373241 A1) and Diggins (US 20160217544 A1), do not teach the cited limitations of “the electronic device of claim 1, wherein the image is a first image, and wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to: obtain other depth information associated with the visual object at a second moment after a first moment corresponding to the first image, obtain difference values between depth values in the alpha channel of the first image, which are indicated by the depth information, and other depth values respectively corresponding to a second image corresponding to the second moment indicated by the other depth information, based on obtaining the difference values in a reference range, generate the second image having the alpha channel comprising the difference values, and based on obtaining the difference values outside the reference range, generate the second image having the alpha channel comprising the other depth values and the other transparencies.”
Claim 12 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The best arts searched, Siver, etc. (US 20180213255 A1), Rondao, etc. (US 20150117518 A1), Hollis, etc. (US 20030184556 A1), Fleureau, etc. (US 20210385454 A1), Boyce, etc. (US 20190373241 A1) and Diggins (US 20160217544 A1), do not teach the cited limitations of “the electronic device of claim 1, further comprises a sensor configured to detect a motion of a user, wherein the image is a first image which represents the visual object moved based on the motion detected by the sensor, and wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to: obtain the depth information to generate the first image using first sensor data detected from the sensor at a first moment, detect second sensor data from the sensor at a second moment after the first moment, identify a difference between the first sensor data detected at the first moment and the second sensor data detected at the second moment, based on identifying that the difference is within a reference range, generate a second image corresponding to the second moment, wherein the alpha channel of the second image comprises difference values between the depth values included in the alpha channel of the first image, and other depth values indicated by other depth information obtained based on the sensor data at the second moment, and based on identifying that the difference is outside the reference range, generate the second image corresponding to the second moment, wherein the alpha channel of the second image comprises the other depth values.”
Claim 13 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The best arts searched, Siver, etc. (US 20180213255 A1), Rondao, etc. (US 20150117518 A1), Hollis, etc. (US 20030184556 A1), Fleureau, etc. (US 20210385454 A1), Boyce, etc. (US 20190373241 A1) and Diggins (US 20160217544 A1), do not teach the cited limitations of “the electronic device of claim 1, further comprises a display assembly comprising a plurality of displays, wherein the instructions, when executed by the at least one processor individually or collectively, further cause the electronic device to: receive an input to display the image, obtain, based on receiving the input, the depth values in the alpha channel of the image, determine, based on the obtained depth values, a binocular parallax of the alpha channel, based on the binocular parallax, display the image on a first display among the plurality of displays, and based on the binocular parallax, display another image representing the visual object shifted based on the binocular parallax on a second display among the plurality of displays.”
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GORDON G LIU whose telephone number is (571)270-0382. The examiner can normally be reached Monday - Friday 8:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Devona E Faulk can be reached at 571-272-7515. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/GORDON G LIU/Primary Examiner, Art Unit 2618