MLchartDataset catalogue

Patent · US2016154999A1 · A1 · US

Objection recognition in a 3d scene

(11) Publication number
US2016154999A1
(21) Application number
14/948,260
(22) Filing date
2015-11-21
(30) Priority date
2014-12-02
(43) Publication date
2016-06-02
(51) IPC
G06K 9/00; G06K 9/46; G06T 7/00; G06T 7/11; G06T 7/136; G06T 7/155; G06T 7/73
(52) CPC
  • G06T Image data processing or generation, in general: 7/155, 17/00, 19/00, 2207/10028, 2207/20021, 2207/20036, 2207/20076, 2207/20081, 2207/20112, 2207/20116, 7/0081, 7/0083, 7/0091, 7/10, 7/11, 7/136, 7/50, 7/521, 7/75
  • G06K Graphical data reading; presentation of data; record carriers; handling record carriers: 9/00637, 9/4609
  • G06V Image or video recognition or understanding: 20/00, 20/56, 20/64
(73) Assignee
Nokia Technologies Oy
(72) Inventors
Lixin Fan; Pouria BABAHAJIANI
(54) Title
Objection recognition in a 3d scene
(57) Abstract

A method comprising: obtaining a three-dimensional (3D) point cloud about at least one object of interest; detecting ground and/or building objects from 3D point cloud data using an unsupervised segmentation method; removing the ground and/or building objects from the 3D point cloud data; and detecting one or more vertical objects from the remaining 3D point cloud data using a supervised segmentation method.

Full text
View on Google Patents

Claims (1)

  1. A method, comprising: obtaining a three-dimensional (3D) point cloud about at least one object of interest; detecting at least one of ground and building objects from 3D point cloud data using an unsupervised segmentation method; removing the at least one of ground and building objects from the 3D point cloud data; and detecting one or more vertical objects from remaining 3D point cloud data using a supervised segmentation method. 2. The method according to claim 1, further comprising: dividing the 3D cloud point data into rectangular tiles in a horizontal plane; and determining an estimation of a ground plane within the rectangular tiles. 3. The method according to claim 2, wherein determining the estimation of the ground plane within the rectangular tiles comprises: dividing the rectangular tiles into a plurality of grid cells; searching a minimal-z-value point within each grid cell; searching points in the each grid cell having a z-value within a first predetermined threshold from the minimal-z-value of a grid cell; collecting the points having the z-value within the first predetermined threshold from the minimal-z-value of the grid cell from each grid cell; estimating the ground plane of a tile on the basis of the collected points; and determining points locating within a second predetermined threshold from the estimated ground plane to comprise ground points of the tile. 4. The method according to claim 1, further comprising: projecting 3D points to range image pixels in a horizontal plane; defining values of the range image pixels as a function of number of 3D points projected to the range image pixels and a maximal z-value among the 3D points projected to the range image pixels; defining a geodesic elongation value for objects detected in the range image pixels; and distinguishing buildings from other objects on the basis of the geodesic elongation value. 5. The method according to claim 4, the method further comprising: binarizing the values of the range image pixels; applying morphological operations for merging neighboring points in the binarized values of the range image pixels; and extracting contours for finding boundaries of the objects. 6. The method according to claim 1 further comprising applying a voxel based segmentation to the remaining 3D point cloud data, wherein the voxel based segmentation comprises: performing voxilisation of the 3D point cloud data; and merging of voxels into super-voxels and carrying out supervised classification based on discriminative features extracted from the super-voxels. 7. The method according to claim 6 further comprising: merging the 3D points into voxels comprising a plurality of 3D points such that for a selected 3D point, all neighboring 3D points within a third predefined threshold from the selected 3D point are merged into a voxel without exceeding a maximum number of 3D points in the voxel. 8. The method according to claim 6, further comprising: merging any number of the voxels into the super-voxels such that a criteria is fulfilled, the criteria comprising: a minimal geometrical distance between two voxels is smaller than a fourth predefined threshold; and an angle between normal vectors of the two voxels is smaller than a fifth predefined threshold. 9. The method according to claim 6 further comprising: extracting features from the super-voxels for classifying the super-voxels into objects. 10. The method according to claim 9, wherein the features include one or more of: geometrical shape; height of a voxel above ground; horizontal distance of the voxel to a center line of a street; surface normal of the voxel; voxel planarity; density of 3D points in the voxel; and intensity of the voxel. 11. An apparatus, comprising at least one processor, memory including computer program code, the memory and the computer program code configured to, with the at least one processor, cause the apparatus to at least: obtain a three-dimensional (3D) point cloud about at least one object of interest; detect at least one of ground and building objects from 3D point cloud data using an unsupervised segmentation method; remove the at least one of ground and building objects from the 3D point cloud data; and detect one or more vertical objects from remaining 3D point cloud data using a supervised segmentation method. 12. The apparatus according to claim 11, wherein the apparatus is further caused to: divide the 3D cloud point data into rectangular tiles in a horizontal plane; and determine an estimation of a ground plane within the rectangular tiles. 13. The apparatus according to claim 12, wherein to determine the estimation of the ground plane within the rectangular tiles, the apparatus is further caused to: divide the rectangular tiles into a plurality of grid cells; search a minimal-z-value point within each grid cell; search points in the each grid cell having a z-value within a first predetermined threshold from the minimal-z-value of a grid cell; collecting the points having the z-value within the first predetermined threshold from the minimal-z-value of the grid cell from the each grid cell; estimating the ground plane of a tile on the basis of the collected points; and determining points locating within a second predetermined threshold from the estimated ground plane to comprise ground points of the tile. 14. The apparatus according claim 11, wherein the apparatus is further caused to: project 3D points to range image pixels in a horizontal plane; define values of the range image pixels as a function of number of 3D points projected to the range image pixels and a maximal z-value among the 3D points projected to the range image pixels; define a geodesic elongation value for objects detected in the range image pixels; and distinguish buildings from other objects on the basis of the geodesic elongation value. 15. The apparatus according to claim 14, wherein the apparatus is further caused to: binarize the values of the range image pixels; apply morphological operations for merging neighboring points in the binarized values of the range image pixels; and extract contours for finding boundaries of the objects. 16. The apparatus according to claims 11, wherein the apparatus is further caused to apply a voxel based segmentation to the remaining 3D point cloud data, and wherein to apply the voxel based segmentation, the apparatus is further caused to: perform voxilisation of the 3D point cloud data; merge voxels into super-voxels; and carry out supervised classification based on discriminative features extracted from the super-voxels. 17. The apparatus according to claim 16, wherein the apparatus is further caused to merge the 3D points into voxels comprising a plurality of 3D points such that for a selected 3D point, all neighboring 3D points within a third predefined threshold from the selected 3D point are merged into a voxel without exceeding a maximum number of 3D points in the voxel. 18. The apparatus according to claim 16, wherein the apparatus is further caused to merging any number of the voxels into the super-voxels such that following criteria is fulfilled: a minimal geometrical distance between two voxels is smaller than a fourth predefined threshold; and an angle between normal vectors of the two voxels is smaller than a fifth predefined threshold. 19. The apparatus according to claim 16, wherein the apparatus is the apparatus is further caused to extract features from the super-voxels for classifying the super-voxels into objects. 20. The apparatus according to claim 19, wherein the features include one or more of: geometrical shape; height of a voxel above ground; horizontal distance of the voxel to a center line of a street; surface normal of the voxel; voxel planarity; density of 3D points in the voxel; and intensity of the voxel. 21. A computer readable storage medium stored with code thereon for use by an apparatus, which when executed by a processor, causes the apparatus to perform: obtaining a three-dimensional (3D) point cloud about at least one object of interest; detecting at least one of ground and building objects from 3D point cloud data using an unsupervised segmentation method; removing the at least one of ground and building objects from the 3D point cloud data; and detecting one or more vertical objects from remaining 3D point cloud data using a supervised segmentation method.

Description

The present invention relates to image processing, and more particularly to a process of objection recognition in 3D street scene.

Automatic urban scene object recognition refers to the process of segmentation and classifying of objects of interest in an image into predefined semantic labels, such as “building”, “tree” or “road”. This typically involves a fixed number of object categories, each of which requires a training model for classifying image segments. While many techniques for two-dimensional (2D) object recognition have been proposed, the accuracy of these systems is to some extent unsatisfactory, because 2D image cues are sensitive to varying imaging conditions such as lighting, shadow etc.

Three-dimensional (3D) object recognition systems using laser scanning, such as Light Detection And Ranging (LiDAR), provide an output of 3D point clouds. 3D point clouds can be used for a number of applications, such as rendering appealing visual effect based on the physical properties of 3D structures and cleaning of raw input 3D point clouds e.g. by removing moving objects (car, bike, person). Other 3D object recognition applications include robotics, intelligent vehicle systems, augmented reality, transportation maps and geological surveys where high resolution digital elevation maps help in detecting subtle topographic features.

However, identifying and recognizing objects despite appearance variation (change in e.g. texture, color or illumination) has turned out to be a surprisingly difficult task for computer vision systems.

Citations (1)

  • US20130096886A1
Record as JSON
{
  "publication_number": "US2016154999A1",
  "country": "US",
  "kind": "A1",
  "title": "Objection recognition in a 3d scene",
  "abstract": "A method comprising: obtaining a three-dimensional (3D) point cloud about at least one object of interest; detecting ground and/or building objects from 3D point cloud data using an unsupervised segmentation method; removing the ground and/or building objects from the 3D point cloud data; and detecting one or more vertical objects from the remaining 3D point cloud data using a supervised segmentation method.",
  "claims": [
    "1. A method, comprising: obtaining a three-dimensional (3D) point cloud about at least one object of interest; detecting at least one of ground and building objects from 3D point cloud data using an unsupervised segmentation method; removing the at least one of ground and building objects from the 3D point cloud data; and detecting one or more vertical objects from remaining 3D point cloud data using a supervised segmentation method. 2. The method according to claim 1, further comprising: dividing the 3D cloud point data into rectangular tiles in a horizontal plane; and determining an estimation of a ground plane within the rectangular tiles. 3. The method according to claim 2, wherein determining the estimation of the ground plane within the rectangular tiles comprises: dividing the rectangular tiles into a plurality of grid cells; searching a minimal-z-value point within each grid cell; searching points in the each grid cell having a z-value within a first predetermined threshold from the minimal-z-value of a grid cell; collecting the points having the z-value within the first predetermined threshold from the minimal-z-value of the grid cell from each grid cell; estimating the ground plane of a tile on the basis of the collected points; and determining points locating within a second predetermined threshold from the estimated ground plane to comprise ground points of the tile. 4. The method according to claim 1, further comprising: projecting 3D points to range image pixels in a horizontal plane; defining values of the range image pixels as a function of number of 3D points projected to the range image pixels and a maximal z-value among the 3D points projected to the range image pixels; defining a geodesic elongation value for objects detected in the range image pixels; and distinguishing buildings from other objects on the basis of the geodesic elongation value. 5. The method according to claim 4, the method further comprising: binarizing the values of the range image pixels; applying morphological operations for merging neighboring points in the binarized values of the range image pixels; and extracting contours for finding boundaries of the objects. 6. The method according to claim 1 further comprising applying a voxel based segmentation to the remaining 3D point cloud data, wherein the voxel based segmentation comprises: performing voxilisation of the 3D point cloud data; and merging of voxels into super-voxels and carrying out supervised classification based on discriminative features extracted from the super-voxels. 7. The method according to claim 6 further comprising: merging the 3D points into voxels comprising a plurality of 3D points such that for a selected 3D point, all neighboring 3D points within a third predefined threshold from the selected 3D point are merged into a voxel without exceeding a maximum number of 3D points in the voxel. 8. The method according to claim 6, further comprising: merging any number of the voxels into the super-voxels such that a criteria is fulfilled, the criteria comprising: a minimal geometrical distance between two voxels is smaller than a fourth predefined threshold; and an angle between normal vectors of the two voxels is smaller than a fifth predefined threshold. 9. The method according to claim 6 further comprising: extracting features from the super-voxels for classifying the super-voxels into objects. 10. The method according to claim 9, wherein the features include one or more of: geometrical shape; height of a voxel above ground; horizontal distance of the voxel to a center line of a street; surface normal of the voxel; voxel planarity; density of 3D points in the voxel; and intensity of the voxel. 11. An apparatus, comprising at least one processor, memory including computer program code, the memory and the computer program code configured to, with the at least one processor, cause the apparatus to at least: obtain a three-dimensional (3D) point cloud about at least one object of interest; detect at least one of ground and building objects from 3D point cloud data using an unsupervised segmentation method; remove the at least one of ground and building objects from the 3D point cloud data; and detect one or more vertical objects from remaining 3D point cloud data using a supervised segmentation method. 12. The apparatus according to claim 11, wherein the apparatus is further caused to: divide the 3D cloud point data into rectangular tiles in a horizontal plane; and determine an estimation of a ground plane within the rectangular tiles. 13. The apparatus according to claim 12, wherein to determine the estimation of the ground plane within the rectangular tiles, the apparatus is further caused to: divide the rectangular tiles into a plurality of grid cells; search a minimal-z-value point within each grid cell; search points in the each grid cell having a z-value within a first predetermined threshold from the minimal-z-value of a grid cell; collecting the points having the z-value within the first predetermined threshold from the minimal-z-value of the grid cell from the each grid cell; estimating the ground plane of a tile on the basis of the collected points; and determining points locating within a second predetermined threshold from the estimated ground plane to comprise ground points of the tile. 14. The apparatus according claim 11, wherein the apparatus is further caused to: project 3D points to range image pixels in a horizontal plane; define values of the range image pixels as a function of number of 3D points projected to the range image pixels and a maximal z-value among the 3D points projected to the range image pixels; define a geodesic elongation value for objects detected in the range image pixels; and distinguish buildings from other objects on the basis of the geodesic elongation value. 15. The apparatus according to claim 14, wherein the apparatus is further caused to: binarize the values of the range image pixels; apply morphological operations for merging neighboring points in the binarized values of the range image pixels; and extract contours for finding boundaries of the objects. 16. The apparatus according to claims 11, wherein the apparatus is further caused to apply a voxel based segmentation to the remaining 3D point cloud data, and wherein to apply the voxel based segmentation, the apparatus is further caused to: perform voxilisation of the 3D point cloud data; merge voxels into super-voxels; and carry out supervised classification based on discriminative features extracted from the super-voxels. 17. The apparatus according to claim 16, wherein the apparatus is further caused to merge the 3D points into voxels comprising a plurality of 3D points such that for a selected 3D point, all neighboring 3D points within a third predefined threshold from the selected 3D point are merged into a voxel without exceeding a maximum number of 3D points in the voxel. 18. The apparatus according to claim 16, wherein the apparatus is further caused to merging any number of the voxels into the super-voxels such that following criteria is fulfilled: a minimal geometrical distance between two voxels is smaller than a fourth predefined threshold; and an angle between normal vectors of the two voxels is smaller than a fifth predefined threshold. 19. The apparatus according to claim 16, wherein the apparatus is the apparatus is further caused to extract features from the super-voxels for classifying the super-voxels into objects. 20. The apparatus according to claim 19, wherein the features include one or more of: geometrical shape; height of a voxel above ground; horizontal distance of the voxel to a center line of a street; surface normal of the voxel; voxel planarity; density of 3D points in the voxel; and intensity of the voxel. 21. A computer readable storage medium stored with code thereon for use by an apparatus, which when executed by a processor, causes the apparatus to perform: obtaining a three-dimensional (3D) point cloud about at least one object of interest; detecting at least one of ground and building objects from 3D point cloud data using an unsupervised segmentation method; removing the at least one of ground and building objects from the 3D point cloud data; and detecting one or more vertical objects from remaining 3D point cloud data using a supervised segmentation method."
  ],
  "description_excerpt": "The present invention relates to image processing, and more particularly to a process of objection recognition in 3D street scene.\n\nAutomatic urban scene object recognition refers to the process of segmentation and classifying of objects of interest in an image into predefined semantic labels, such as “building”, “tree” or “road”. This typically involves a fixed number of object categories, each of which requires a training model for classifying image segments. While many techniques for two-dimensional (2D) object recognition have been proposed, the accuracy of these systems is to some extent unsatisfactory, because 2D image cues are sensitive to varying imaging conditions such as lighting, shadow etc.\n\nThree-dimensional (3D) object recognition systems using laser scanning, such as Light Detection And Ranging (LiDAR), provide an output of 3D point clouds. 3D point clouds can be used for a number of applications, such as rendering appealing visual effect based on the physical properties of 3D structures and cleaning of raw input 3D point clouds e.g. by removing moving objects (car, bike, person). Other 3D object recognition applications include robotics, intelligent vehicle systems, augmented reality, transportation maps and geological surveys where high resolution digital elevation maps help in detecting subtle topographic features.\n\nHowever, identifying and recognizing objects despite appearance variation (change in e.g. texture, color or illumination) has turned out to be a surprisingly difficult task for computer vision systems.",
  "cpc": [
    "G06T 7/155",
    "G06K 9/00637",
    "G06K 9/4609",
    "G06T 17/00",
    "G06T 19/00",
    "G06T 2207/10028",
    "G06T 2207/20021",
    "G06T 2207/20036",
    "G06T 2207/20076",
    "G06T 2207/20081",
    "G06T 2207/20112",
    "G06T 2207/20116",
    "G06T 7/0081",
    "G06T 7/0083",
    "G06T 7/0091",
    "G06T 7/10",
    "G06T 7/11",
    "G06T 7/136",
    "G06T 7/50",
    "G06T 7/521",
    "G06T 7/75",
    "G06V 20/00",
    "G06V 20/56",
    "G06V 20/64"
  ],
  "ipc": [
    "G06K 9/00",
    "G06K 9/46",
    "G06T 7/00",
    "G06T 7/11",
    "G06T 7/136",
    "G06T 7/155",
    "G06T 7/73"
  ],
  "assignees": [
    "Nokia Technologies Oy"
  ],
  "inventors": [
    "Lixin Fan",
    "Pouria BABAHAJIANI"
  ],
  "filing_date": "2015-11-21",
  "publication_date": "2016-06-02",
  "priority_date": "2014-12-02",
  "application_number": "US-201514948260-A",
  "family_id": "52349758",
  "cited_by_count": 186,
  "citations": [
    "US20130096886A1"
  ]
}

Record 4,802 of 8,000 in Patents full text (MLC-0201). Request the full dataset.