There are 12 ship types in the inventory, known as vessel classifications. These are imported from a CSV provided by PLA. A sample is shown below.
# A tibble: 67 x 6registry_classaggregated_classgroupcoderelease_heightfuel_type<chr><chr><dbl><chr><dbl><chr>1ferry Passenger (ferry) 1OFY5Diesel2cuttersuctiondredgerDredger2DCS17.5Diesel3dredgerDredger2DDR17.5Diesel4hopperdredgerDredger2DHD17.5Diesel5suctionhopperdredgerDredger2DSH17.5Diesel6sandsuctiondredgerDredger2DSS17.5Diesel7trailingsuctiondredgerDredger2DTD17.5Diesel8trailingsuctionhopperdredgerDredger2DTS17.5Diesel9 fishing (general) Fishing2FFS17.5Diesel10 trawler (Alltypes) Fishing2FTR17.5DieselEach type of ship has a code and group. There are four groups.
PLA 2016 emissions covering London are imported as a CSV. These contain for each pollutant, ship type and cellid (LAEI exact cut) the estimated shipping emissions in kg per annum for that area (approx. 1km x 1km). A sample is shown below.
# A tibble: 3,159 x 5ship_typepollutantcellidsailingberth<chr><chr><dbl><dbl><dbl>1BulkcarrierNOx158441.802BulkcarrierNOx91122.603BulkcarrierNOx9070.504BulkcarrierNOx22119.5305BulkcarrierNOx220549.806BulkcarrierNOx224610.507BulkcarrierNOx86072.908BulkcarrierNOx220868.509BulkcarrierNOx22122.99010BulkcarrierNOx158740.70The vessel_classifications are joined to the inventory_export_2016.csv by ship_type. However there are three ship_type's in the emissions that do not exactly match with the vessel_classifications file. These are edited to force them to link, as below.
| Emissions ship_type | Vessel classification ship_type |
|---|---|
| 'RoRo Cargo / Vehicle' | 'RoRo Cargo/Vehicle' |
| 'Cruise ship' | 'Passenger (cruise)' |
| 'Passenger' | 'Passenger (ferry)' |
The emissions now look like this:
# A tibble: 3,159 x 6ship_typepollutantcellidsailingberthgroup<chr><chr><dbl><dbl><dbl><dbl>1BulkcarrierNOx158441.8032BulkcarrierNOx91122.6033BulkcarrierNOx9070.5034BulkcarrierNOx22119.53035BulkcarrierNOx220549.8036BulkcarrierNOx224610.5037BulkcarrierNOx86072.903Using the group, pollutant and cellid columns they are now grouped to give the below.
# A tibble: 1,590 x 5pollutantcellidgroupsailingberth<chr><dbl><dbl><dbl><dbl>1NOx231154.602NOx23126778.2.093NOx23131429.04NOx23143970.4567.5NOx23213.3506NOx2322521.1.547NOx232395.708NOx232410367.29460.9NOx23412.28010NOx2342332.0This data is now joined to a geopackage of the LAEI grid exact cuts using the cellid identifier. Then the LAEI exact cuts are dissolved to create regular 1km by 1km grids n.b. this is not a spatial aggregate, actually the grid centroid coordinates are used to create the grids, and then linked. The data (grid_emissions) now looks like this:
Simplefeaturecollectionwith933featuresand5fieldsgeometrytype:POLYGONdimension:XYbbox:xmin:523000ymin:174000xmax:558000ymax:187000
epsg (SRID):27700proj4string:+proj=tmerc+lat_0=49+lon_0=-2+k=0.9996012717+x_0=400000+y_0=-100000+ellps=airy+towgs84=446.448,-125.157,542.06,0.15,0.247,0.842,-20.489+units=m+no_defs# A tibble: 933 x 6pollutantgrouplarge_grid_idsailingberthgeometry<chr><dbl><dbl><dbl><dbl><POLYGON [m]>1NOx195590.170 ((548000183000, 548000182000, 547000182000, 547000183000, 548000183000))
2NOx197170.260 ((534000182000, 534000181000, 533000181000, 533000182000, 534000182000))
3NOx1972820.90 ((545000182000, 545000181000, 544000181000, 544000182000, 545000182000))
4NOx1972982.80.01 ((546000182000, 546000181000, 545000181000, 545000182000, 546000182000))
5NOx1973062.90 ((547000182000, 547000181000, 546000181000, 546000182000, 547000182000))
6NOx1973159.40 ((548000182000, 548000181000, 547000181000, 547000182000, 548000182000))
7NOx19732108.0 ((549000182000, 549000181000, 548000181000, 548000182000, 549000182000))
8NOx1973345.40 ((550000182000, 550000181000, 549000181000, 549000182000, 550000182000))
9NOx197349.80 ((551000182000, 551000181000, 550000181000, 550000182000, 551000182000))
10NOx1988622179.34.1 ((531000181000, 531000180000, 530000180000, 530000181000, 531000181000))
# ... with 923 more rowsA 20m x 20m polygon grid (small_grid) covering the extent of the grid_emissions was now created. This was then intersected and clipped with a unique geographical representation of the emissions grid i.e. just one cell for each emission area, rather than twelve cells (3 pollutants x 4 groups). The link between the small_grid and the grid_emissions is by large_grid_id.The small_grid data looks like this:
Simplefeaturecollectionwith334932featuresand2fieldsgeometrytype:GEOMETRYdimension:XYbbox:xmin:523000ymin:174000xmax:558000ymax:187000
epsg (SRID):27700proj4string:+proj=tmerc+lat_0=49+lon_0=-2+k=0.9996012717+x_0=400000+y_0=-100000+ellps=airy+towgs84=446.448,-125.157,542.06,0.15,0.247,0.842,-20.489+units=m+no_defsFirst10features:large_grid_idsmall_grid_idgeometry195591 POINT (547000182000)
295592 LINESTRING (547020182000, ...395593 LINESTRING (547040182000, ...495594 LINESTRING (547060182000, ...595595 LINESTRING (547080182000, ...695596 LINESTRING (547100182000, ...795597 LINESTRING (547120182000, ...895598 LINESTRING (547140182000, ...995599 LINESTRING (547160182000, ...10955910 LINESTRING (547180182000, ...366 days of AIS data for 2016 were provided by the PLA. These were processed using the Python libais library to extract the latitude, longitude and ship_type (group), and the data exported as an R Dataframe. A sample is shown below.
lonlatVESSEL_TYPE10.396855051.44474XTG20.341360051.46286GPC30.326833351.45400GGC40.252336751.46938URR50.396293351.44480XTG60.349478351.45776GPCOn 1 January 2016 (taken as an example day) there were 1,368,454 GPS points in the dataset, 336,967 (25%) of which had no category. The remaining categories and counts are as below.
BBUBCEBWCDTDDTSGGCGPCOFYOSUOYTPRRRSRTCOTEOTPDURRXFFXTG1122772189597107438333207300641703411652443235161902667953785822911781040554413542119259246The AIS data was imported in turn, and spatially joined to the small_grid to create the small_grid_result. This being that each small_grid contained the count of the total number of GPS points, per group, that had been recorded in that grid square over 2016. 20m cells with less than 20 GPS points in them were treated as GPS drift error and were discarded. The map and data below show the annual count of GPS points within each grid square.
Simplefeaturecollectionwith1223680featuresand5fieldsgeometrytype:GEOMETRYdimension:XYbbox:xmin:523000ymin:174000xmax:558000ymax:187000
epsg (SRID):27700proj4string:+proj=tmerc+lat_0=49+lon_0=-2+k=0.9996012717+x_0=400000+y_0=-100000+ellps=airy+towgs84=446.448,-125.157,542.06,0.15,0.247,0.842,-20.489+units=m+no_defsFirst10features:cellidsmall_grid_idgroupcountberth_namegeometry123111NA<NA> POLYGON ((555000177020, 55...223121NA<NA> POLYGON ((555000177000, 55...323131NA<NA> POLYGON ((555020177000, 55...423141NA<NA> POLYGON ((555040177000, 55...523151NA<NA> POLYGON ((555060177000, 55...623161NA<NA> POLYGON ((555080177000, 55...723171NA<NA> POLYGON ((555100177000, 55...823181NA<NA> POLYGON ((555120177000, 55...923191NA<NA> POLYGON ((555140177000, 55...10231101NA<NA> POLYGON ((555160177000, 55...The grid_emissions are split between sailing and berth. The berth emissions will be distributed to the berths that are within each large_grid_id, weighted by the number of GPS points within the small_grid cell of the berth. A berths.shp was therefore imported, and joined to the small_grid, adding an extra column to identify if that cell held a berth or not. The berths data looked like this
Simplefeaturecollectionwith180featuresand1fieldgeometrytype:POINTdimension:XYbbox:xmin:-1.797693e+308ymin:-1.797693e+308xmax:589000ymax:183040
epsg (SRID):27700proj4string:+proj=tmerc+lat_0=49+lon_0=-2+k=0.9996012717+x_0=400000+y_0=-100000+ellps=airy+towgs84=446.448,-125.157,542.06,0.15,0.247,0.842,-20.489+units=m+no_defsFirst10features:berth_namegeometry1TowerStairsTier POINT (533225180490)
2StKatharine's Dock POINT (533875 180325)3 President Quay POINT (533970 180200)4 Express Wharf POINT (537000 179700)5 West India Lock POINT (538320 179920)6 Pura Foods POINT (539060 180660)7 Orchard Wharf POINT (539230 180710)8 Thames Wharf POINT (539600 180500)9 Royal Primrose Works POINT (540200 179850)10 Silvertown Wharf POINT (540440 179720)The small_grid_result was grouped by large_grid_id to calculate the total number of GPS points in the parent grids. Then the count of GPS points in each 20m by 20m grid, was divided by the total of GPS points within the larger grid (1km by 1km) for that area, to give a weighting for each small grid to draw down the emissions. The grid cells containing berths were calculated in a similar manner but indenpendantly i.e. only cells that contained berths were divided by the berth_count. This created the contribution column.
Simplefeaturecollectionwith126714featuresand8fieldsgeometrytype:POLYGONdimension:XYbbox:xmin:523860ymin:174420xmax:558000ymax:186900
epsg (SRID):27700proj4string:+proj=tmerc+lat_0=49+lon_0=-2+k=0.9996012717+x_0=400000+y_0=-100000+ellps=airy+towgs84=446.448,-125.157,542.06,0.15,0.247,0.842,-20.489+units=m+no_defsFirst10features:large_grid_idsmall_grid_idgroupcountberth_namesailing_countberth_countcontributiongeometry19717292812<NA>9NA0.222222222 POLYGON ((533280181060, 53...29717298113<NA>9NA0.333333333 POLYGON ((533300181080, 53...39717308512<NA>9NA0.222222222 POLYGON ((533300181120, 53...49717339711<NA>9NA0.111111111 POLYGON ((533300181240, 53...59717364211<NA>9NA0.111111111 POLYGON ((533000181360, 53...69728549612<NA>719NA0.002781641 POLYGON ((544700181000, 54...79728549715<NA>719NA0.006954103 POLYGON ((544720181000, 54...89728549812<NA>719NA0.002781641 POLYGON ((544740181000, 54...99728549915<NA>719NA0.006954103 POLYGON ((544760181000, 54...109728550014<NA>719NA0.005563282 POLYGON ((544780181000, 54...The small_grid was now duplicated 3 times, a pollutant column added to each and populated (PM2.5, NOx and PM), and then joined back together again. It was then joined to the emissions by pollutant, group and unique_geom_id. Subsequently the contribution was multiplied by the emission for that large grid square, and stored as emissions.
There are 1km by 1km emission grid areas from the PLA, with emissions for sailing/berth, where there is no GPS data.
small_grid_result %>% as_tibble() %>% group_by(large_grid_id, pollutant, group) %>% summarise(gps_count= sum(count, na.rm=T)) %>% left_join(grid_emissions, ., by= c("large_grid_id"="large_grid_id",
"group"="group",
"pollutant"="pollutant")) %>% filter(gps_count==0& (sailing>0|berth>0))
# A tibble: 62 x 6pollutantgrouplarge_grid_idsailingberthgps_count<chr><dbl><dbl><dbl><dbl><int>1NOx195590.17002NOx198930.02003NOx1100810.01004NOx293830.16005NOx293840.900006NOx295561.55007NOx295575.760.4808NOx295595.1712.909NOx295600.640010NOx297170.0100# ... with 52 more rowsHowever the emissions in these large_grid_squares are a very small proportion of the total emissions (shown below) so we have decided not to investigate this issue further.
| Pollutant | Sailing | Berth |
|---|---|---|
| NOx | 0.003% | 0.006% |
| PM | 0.002% | 0.006% |
| PM2.5 | 0.002% | 0.006% |
Due to the way that the emissions were calculated on a grid, there are some artefacts when visualising the data, such as shown below. These have not been resolved (as there is no easy way to do so). It is a limitation of the underlying data/method before our work.









