Repository files navigation

Build and Test

Upload Python Package

Known Vulnerabilities

SparkAutoMapper

Fluent API to map data from one view to another in Spark.

Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Since this is just Python, you can use any Python editor. Since everything is typed using Python typings, most editors will auto-complete and warn you when you do something wrong

Usage

pip install sparkautomapper

Documentation

https://icanbwell.github.io/SparkAutoMapper/

SparkAutoMapper input and output

You can pass either a dataframe to SparkAutoMapper or specify the name of a Spark view to read from.

You can receive the result as a dataframe or (optionally) pass in the name of a view where you want the result.

Dynamic Typing Examples

Set a column in destination to a text value (read from pass in data frame and return the result in a new dataframe)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to a text value (read from a Spark view and put result in another Spark view)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to an int value

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=1050
)

Copy a column (src1) from source_view to destination view column (dst1)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column("src1")
)

Or you can use the shortcut for specifying a column (wrap column name in [])

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]"
)

Convert data type for a column (or string literal)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
birthDate=A.date(A.column("date_of_birth"))
)

Use a Spark SQL Expression (Any valid Spark SQL expression can be used)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Specify multiple transformations

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]",
birthDate=A.date("[date_of_birth]"),
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Use variables or parameters

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)

Use conditional logic

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)
ifparameters["customer"] =="Microsoft":
mapper=mapper.columns(
important_customer=1,
customer_name=parameters["customer"]
)
returnmapper

Using nested array columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).withColumn(
dst2=A.list(
[
"address1",
"address2"
]
)
)

Using nested struct columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.complex(
use="usual",
family="imran"
)
)

Using lists of structs

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.list(
[
A.complex(
use="usual",
family="imran"
),
A.complex(
use="usual",
family="[last_name]"
)
]
)
)

Executing the AutoMapper

spark.createDataFrame(
[
(1, 'Qureshi', 'Imran'),
(2, 'Vidal', 'Michael'),
],
['member_id', 'last_name', 'first_name']
).createOrReplaceTempView("patients")
source_df: DataFrame=spark.table("patients")
df=source_df.select("member_id")
df.createOrReplaceTempView("members")
result_df: DataFrame=mapper.transform(df=df)

Statically Typed Examples

To improve the auto-complete and syntax checking even more, you can define Complex types:

Define a custom data type:

fromspark_auto_mapper.type_definitions.automapper_defined_typesimportAutoMapperTextInputTypefromspark_auto_mapper.helpers.automapper_value_parserimportAutoMapperValueParserfromspark_auto_mapper.data_types.dateimportAutoMapperDateDataTypefromspark_auto_mapper.data_types.listimportAutoMapperListfromspark_auto_mapper_fhir.fhir_types.automapper_fhir_data_type_complex_baseimportAutoMapperFhirDataTypeComplexBaseclassAutoMapperFhirDataTypePatient(AutoMapperFhirDataTypeComplexBase):
# noinspection PyPep8Namingdef__init__(self,
id_: AutoMapperTextInputType,
birthDate: AutoMapperDateDataType,
name: AutoMapperList,
gender: AutoMapperTextInputType
) ->None:
super().__init__()
self.value=dict(
id=AutoMapperValueParser.parse_value(id_),
birthDate=AutoMapperValueParser.parse_value(birthDate),
name=AutoMapperValueParser.parse_value(name),
gender=AutoMapperValueParser.parse_value(gender)
)

Now you get auto-complete and syntax checking:

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapperFhir(
view="members",
source_view="patients",
keys=["member_id"]
).withResource(
resource=F.patient(
id_=A.column("a.member_id"),
birthDate=A.date(
A.column("date_of_birth")
),
name=A.list(
F.human_name(
use="usual",
family=A.column("last_name")
)
),
gender="female"
)
)

Publishing a new package

  1. Edit VERSION to increment the version
  2. Create a new release
  3. The GitHub Action should automatically kick in and publish the package
  4. You can see the status in the Actions tab

About

Fluent API to map data from one view to another in Spark. Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Resources

Contributing

Stars

7 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all \u003cpre\u003e\u003ccode\u003e blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks"); } } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); } })(); (function(){ try { var __m = "github.com"; var __re = new RegExp('^' + "github\\.com" + '
Skip to content

Repository files navigation

Build and Test

Upload Python Package

Known Vulnerabilities

SparkAutoMapper

Fluent API to map data from one view to another in Spark.

Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Since this is just Python, you can use any Python editor. Since everything is typed using Python typings, most editors will auto-complete and warn you when you do something wrong

Usage

pip install sparkautomapper

Documentation

https://icanbwell.github.io/SparkAutoMapper/

SparkAutoMapper input and output

You can pass either a dataframe to SparkAutoMapper or specify the name of a Spark view to read from.

You can receive the result as a dataframe or (optionally) pass in the name of a view where you want the result.

Dynamic Typing Examples

Set a column in destination to a text value (read from pass in data frame and return the result in a new dataframe)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to a text value (read from a Spark view and put result in another Spark view)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to an int value

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=1050
)

Copy a column (src1) from source_view to destination view column (dst1)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column("src1")
)

Or you can use the shortcut for specifying a column (wrap column name in [])

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]"
)

Convert data type for a column (or string literal)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
birthDate=A.date(A.column("date_of_birth"))
)

Use a Spark SQL Expression (Any valid Spark SQL expression can be used)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Specify multiple transformations

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]",
birthDate=A.date("[date_of_birth]"),
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Use variables or parameters

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)

Use conditional logic

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)
ifparameters["customer"] =="Microsoft":
mapper=mapper.columns(
important_customer=1,
customer_name=parameters["customer"]
)
returnmapper

Using nested array columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).withColumn(
dst2=A.list(
[
"address1",
"address2"
]
)
)

Using nested struct columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.complex(
use="usual",
family="imran"
)
)

Using lists of structs

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.list(
[
A.complex(
use="usual",
family="imran"
),
A.complex(
use="usual",
family="[last_name]"
)
]
)
)

Executing the AutoMapper

spark.createDataFrame(
[
(1, 'Qureshi', 'Imran'),
(2, 'Vidal', 'Michael'),
],
['member_id', 'last_name', 'first_name']
).createOrReplaceTempView("patients")
source_df: DataFrame=spark.table("patients")
df=source_df.select("member_id")
df.createOrReplaceTempView("members")
result_df: DataFrame=mapper.transform(df=df)

Statically Typed Examples

To improve the auto-complete and syntax checking even more, you can define Complex types:

Define a custom data type:

fromspark_auto_mapper.type_definitions.automapper_defined_typesimportAutoMapperTextInputTypefromspark_auto_mapper.helpers.automapper_value_parserimportAutoMapperValueParserfromspark_auto_mapper.data_types.dateimportAutoMapperDateDataTypefromspark_auto_mapper.data_types.listimportAutoMapperListfromspark_auto_mapper_fhir.fhir_types.automapper_fhir_data_type_complex_baseimportAutoMapperFhirDataTypeComplexBaseclassAutoMapperFhirDataTypePatient(AutoMapperFhirDataTypeComplexBase):
# noinspection PyPep8Namingdef__init__(self,
id_: AutoMapperTextInputType,
birthDate: AutoMapperDateDataType,
name: AutoMapperList,
gender: AutoMapperTextInputType
) ->None:
super().__init__()
self.value=dict(
id=AutoMapperValueParser.parse_value(id_),
birthDate=AutoMapperValueParser.parse_value(birthDate),
name=AutoMapperValueParser.parse_value(name),
gender=AutoMapperValueParser.parse_value(gender)
)

Now you get auto-complete and syntax checking:

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapperFhir(
view="members",
source_view="patients",
keys=["member_id"]
).withResource(
resource=F.patient(
id_=A.column("a.member_id"),
birthDate=A.date(
A.column("date_of_birth")
),
name=A.list(
F.human_name(
use="usual",
family=A.column("last_name")
)
),
gender="female"
)
)

Publishing a new package

  1. Edit VERSION to increment the version
  2. Create a new release
  3. The GitHub Action should automatically kick in and publish the package
  4. You can see the status in the Actions tab

About

Fluent API to map data from one view to another in Spark. Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Resources

Contributing

Stars

7 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Build and Test

Upload Python Package

Known Vulnerabilities

SparkAutoMapper

Fluent API to map data from one view to another in Spark.

Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Since this is just Python, you can use any Python editor. Since everything is typed using Python typings, most editors will auto-complete and warn you when you do something wrong

Usage

pip install sparkautomapper

Documentation

https://icanbwell.github.io/SparkAutoMapper/

SparkAutoMapper input and output

You can pass either a dataframe to SparkAutoMapper or specify the name of a Spark view to read from.

You can receive the result as a dataframe or (optionally) pass in the name of a view where you want the result.

Dynamic Typing Examples

Set a column in destination to a text value (read from pass in data frame and return the result in a new dataframe)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to a text value (read from a Spark view and put result in another Spark view)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to an int value

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=1050
)

Copy a column (src1) from source_view to destination view column (dst1)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column("src1")
)

Or you can use the shortcut for specifying a column (wrap column name in [])

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]"
)

Convert data type for a column (or string literal)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
birthDate=A.date(A.column("date_of_birth"))
)

Use a Spark SQL Expression (Any valid Spark SQL expression can be used)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Specify multiple transformations

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]",
birthDate=A.date("[date_of_birth]"),
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Use variables or parameters

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)

Use conditional logic

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)
ifparameters["customer"] =="Microsoft":
mapper=mapper.columns(
important_customer=1,
customer_name=parameters["customer"]
)
returnmapper

Using nested array columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).withColumn(
dst2=A.list(
[
"address1",
"address2"
]
)
)

Using nested struct columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.complex(
use="usual",
family="imran"
)
)

Using lists of structs

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.list(
[
A.complex(
use="usual",
family="imran"
),
A.complex(
use="usual",
family="[last_name]"
)
]
)
)

Executing the AutoMapper

spark.createDataFrame(
[
(1, 'Qureshi', 'Imran'),
(2, 'Vidal', 'Michael'),
],
['member_id', 'last_name', 'first_name']
).createOrReplaceTempView("patients")
source_df: DataFrame=spark.table("patients")
df=source_df.select("member_id")
df.createOrReplaceTempView("members")
result_df: DataFrame=mapper.transform(df=df)

Statically Typed Examples

To improve the auto-complete and syntax checking even more, you can define Complex types:

Define a custom data type:

fromspark_auto_mapper.type_definitions.automapper_defined_typesimportAutoMapperTextInputTypefromspark_auto_mapper.helpers.automapper_value_parserimportAutoMapperValueParserfromspark_auto_mapper.data_types.dateimportAutoMapperDateDataTypefromspark_auto_mapper.data_types.listimportAutoMapperListfromspark_auto_mapper_fhir.fhir_types.automapper_fhir_data_type_complex_baseimportAutoMapperFhirDataTypeComplexBaseclassAutoMapperFhirDataTypePatient(AutoMapperFhirDataTypeComplexBase):
# noinspection PyPep8Namingdef__init__(self,
id_: AutoMapperTextInputType,
birthDate: AutoMapperDateDataType,
name: AutoMapperList,
gender: AutoMapperTextInputType
) ->None:
super().__init__()
self.value=dict(
id=AutoMapperValueParser.parse_value(id_),
birthDate=AutoMapperValueParser.parse_value(birthDate),
name=AutoMapperValueParser.parse_value(name),
gender=AutoMapperValueParser.parse_value(gender)
)

Now you get auto-complete and syntax checking:

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapperFhir(
view="members",
source_view="patients",
keys=["member_id"]
).withResource(
resource=F.patient(
id_=A.column("a.member_id"),
birthDate=A.date(
A.column("date_of_birth")
),
name=A.list(
F.human_name(
use="usual",
family=A.column("last_name")
)
),
gender="female"
)
)

Publishing a new package

  1. Edit VERSION to increment the version
  2. Create a new release
  3. The GitHub Action should automatically kick in and publish the package
  4. You can see the status in the Actions tab

About

Fluent API to map data from one view to another in Spark. Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Resources

Contributing

Stars

7 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length \u003e 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Build and Test

Upload Python Package

Known Vulnerabilities

SparkAutoMapper

Fluent API to map data from one view to another in Spark.

Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Since this is just Python, you can use any Python editor. Since everything is typed using Python typings, most editors will auto-complete and warn you when you do something wrong

Usage

pip install sparkautomapper

Documentation

https://icanbwell.github.io/SparkAutoMapper/

SparkAutoMapper input and output

You can pass either a dataframe to SparkAutoMapper or specify the name of a Spark view to read from.

You can receive the result as a dataframe or (optionally) pass in the name of a view where you want the result.

Dynamic Typing Examples

Set a column in destination to a text value (read from pass in data frame and return the result in a new dataframe)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to a text value (read from a Spark view and put result in another Spark view)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to an int value

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=1050
)

Copy a column (src1) from source_view to destination view column (dst1)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column("src1")
)

Or you can use the shortcut for specifying a column (wrap column name in [])

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]"
)

Convert data type for a column (or string literal)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
birthDate=A.date(A.column("date_of_birth"))
)

Use a Spark SQL Expression (Any valid Spark SQL expression can be used)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Specify multiple transformations

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]",
birthDate=A.date("[date_of_birth]"),
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Use variables or parameters

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)

Use conditional logic

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)
ifparameters["customer"] =="Microsoft":
mapper=mapper.columns(
important_customer=1,
customer_name=parameters["customer"]
)
returnmapper

Using nested array columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).withColumn(
dst2=A.list(
[
"address1",
"address2"
]
)
)

Using nested struct columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.complex(
use="usual",
family="imran"
)
)

Using lists of structs

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.list(
[
A.complex(
use="usual",
family="imran"
),
A.complex(
use="usual",
family="[last_name]"
)
]
)
)

Executing the AutoMapper

spark.createDataFrame(
[
(1, 'Qureshi', 'Imran'),
(2, 'Vidal', 'Michael'),
],
['member_id', 'last_name', 'first_name']
).createOrReplaceTempView("patients")
source_df: DataFrame=spark.table("patients")
df=source_df.select("member_id")
df.createOrReplaceTempView("members")
result_df: DataFrame=mapper.transform(df=df)

Statically Typed Examples

To improve the auto-complete and syntax checking even more, you can define Complex types:

Define a custom data type:

fromspark_auto_mapper.type_definitions.automapper_defined_typesimportAutoMapperTextInputTypefromspark_auto_mapper.helpers.automapper_value_parserimportAutoMapperValueParserfromspark_auto_mapper.data_types.dateimportAutoMapperDateDataTypefromspark_auto_mapper.data_types.listimportAutoMapperListfromspark_auto_mapper_fhir.fhir_types.automapper_fhir_data_type_complex_baseimportAutoMapperFhirDataTypeComplexBaseclassAutoMapperFhirDataTypePatient(AutoMapperFhirDataTypeComplexBase):
# noinspection PyPep8Namingdef__init__(self,
id_: AutoMapperTextInputType,
birthDate: AutoMapperDateDataType,
name: AutoMapperList,
gender: AutoMapperTextInputType
) ->None:
super().__init__()
self.value=dict(
id=AutoMapperValueParser.parse_value(id_),
birthDate=AutoMapperValueParser.parse_value(birthDate),
name=AutoMapperValueParser.parse_value(name),
gender=AutoMapperValueParser.parse_value(gender)
)

Now you get auto-complete and syntax checking:

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapperFhir(
view="members",
source_view="patients",
keys=["member_id"]
).withResource(
resource=F.patient(
id_=A.column("a.member_id"),
birthDate=A.date(
A.column("date_of_birth")
),
name=A.list(
F.human_name(
use="usual",
family=A.column("last_name")
)
),
gender="female"
)
)

Publishing a new package

  1. Edit VERSION to increment the version
  2. Create a new release
  3. The GitHub Action should automatically kick in and publish the package
  4. You can see the status in the Actions tab

About

Fluent API to map data from one view to another in Spark. Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Resources

Contributing

Stars

7 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

Repository files navigation

Build and Test

Upload Python Package

Known Vulnerabilities

SparkAutoMapper

Fluent API to map data from one view to another in Spark.

Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Since this is just Python, you can use any Python editor. Since everything is typed using Python typings, most editors will auto-complete and warn you when you do something wrong

Usage

pip install sparkautomapper

Documentation

https://icanbwell.github.io/SparkAutoMapper/

SparkAutoMapper input and output

You can pass either a dataframe to SparkAutoMapper or specify the name of a Spark view to read from.

You can receive the result as a dataframe or (optionally) pass in the name of a view where you want the result.

Dynamic Typing Examples

Set a column in destination to a text value (read from pass in data frame and return the result in a new dataframe)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to a text value (read from a Spark view and put result in another Spark view)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to an int value

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=1050
)

Copy a column (src1) from source_view to destination view column (dst1)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column("src1")
)

Or you can use the shortcut for specifying a column (wrap column name in [])

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]"
)

Convert data type for a column (or string literal)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
birthDate=A.date(A.column("date_of_birth"))
)

Use a Spark SQL Expression (Any valid Spark SQL expression can be used)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Specify multiple transformations

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]",
birthDate=A.date("[date_of_birth]"),
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Use variables or parameters

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)

Use conditional logic

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)
ifparameters["customer"] =="Microsoft":
mapper=mapper.columns(
important_customer=1,
customer_name=parameters["customer"]
)
returnmapper

Using nested array columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).withColumn(
dst2=A.list(
[
"address1",
"address2"
]
)
)

Using nested struct columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.complex(
use="usual",
family="imran"
)
)

Using lists of structs

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.list(
[
A.complex(
use="usual",
family="imran"
),
A.complex(
use="usual",
family="[last_name]"
)
]
)
)

Executing the AutoMapper

spark.createDataFrame(
[
(1, 'Qureshi', 'Imran'),
(2, 'Vidal', 'Michael'),
],
['member_id', 'last_name', 'first_name']
).createOrReplaceTempView("patients")
source_df: DataFrame=spark.table("patients")
df=source_df.select("member_id")
df.createOrReplaceTempView("members")
result_df: DataFrame=mapper.transform(df=df)

Statically Typed Examples

To improve the auto-complete and syntax checking even more, you can define Complex types:

Define a custom data type:

fromspark_auto_mapper.type_definitions.automapper_defined_typesimportAutoMapperTextInputTypefromspark_auto_mapper.helpers.automapper_value_parserimportAutoMapperValueParserfromspark_auto_mapper.data_types.dateimportAutoMapperDateDataTypefromspark_auto_mapper.data_types.listimportAutoMapperListfromspark_auto_mapper_fhir.fhir_types.automapper_fhir_data_type_complex_baseimportAutoMapperFhirDataTypeComplexBaseclassAutoMapperFhirDataTypePatient(AutoMapperFhirDataTypeComplexBase):
# noinspection PyPep8Namingdef__init__(self,
id_: AutoMapperTextInputType,
birthDate: AutoMapperDateDataType,
name: AutoMapperList,
gender: AutoMapperTextInputType
) ->None:
super().__init__()
self.value=dict(
id=AutoMapperValueParser.parse_value(id_),
birthDate=AutoMapperValueParser.parse_value(birthDate),
name=AutoMapperValueParser.parse_value(name),
gender=AutoMapperValueParser.parse_value(gender)
)

Now you get auto-complete and syntax checking:

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapperFhir(
view="members",
source_view="patients",
keys=["member_id"]
).withResource(
resource=F.patient(
id_=A.column("a.member_id"),
birthDate=A.date(
A.column("date_of_birth")
),
name=A.list(
F.human_name(
use="usual",
family=A.column("last_name")
)
),
gender="female"
)
)

Publishing a new package

  1. Edit VERSION to increment the version
  2. Create a new release
  3. The GitHub Action should automatically kick in and publish the package
  4. You can see the status in the Actions tab

About

Fluent API to map data from one view to another in Spark. Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Resources

Contributing

Stars

7 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Build and Test

Upload Python Package

Known Vulnerabilities

SparkAutoMapper

Fluent API to map data from one view to another in Spark.

Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Since this is just Python, you can use any Python editor. Since everything is typed using Python typings, most editors will auto-complete and warn you when you do something wrong

Usage

pip install sparkautomapper

Documentation

https://icanbwell.github.io/SparkAutoMapper/

SparkAutoMapper input and output

You can pass either a dataframe to SparkAutoMapper or specify the name of a Spark view to read from.

You can receive the result as a dataframe or (optionally) pass in the name of a view where you want the result.

Dynamic Typing Examples

Set a column in destination to a text value (read from pass in data frame and return the result in a new dataframe)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to a text value (read from a Spark view and put result in another Spark view)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to an int value

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=1050
)

Copy a column (src1) from source_view to destination view column (dst1)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column("src1")
)

Or you can use the shortcut for specifying a column (wrap column name in [])

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]"
)

Convert data type for a column (or string literal)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
birthDate=A.date(A.column("date_of_birth"))
)

Use a Spark SQL Expression (Any valid Spark SQL expression can be used)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Specify multiple transformations

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]",
birthDate=A.date("[date_of_birth]"),
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Use variables or parameters

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)

Use conditional logic

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)
ifparameters["customer"] =="Microsoft":
mapper=mapper.columns(
important_customer=1,
customer_name=parameters["customer"]
)
returnmapper

Using nested array columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).withColumn(
dst2=A.list(
[
"address1",
"address2"
]
)
)

Using nested struct columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.complex(
use="usual",
family="imran"
)
)

Using lists of structs

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.list(
[
A.complex(
use="usual",
family="imran"
),
A.complex(
use="usual",
family="[last_name]"
)
]
)
)

Executing the AutoMapper

spark.createDataFrame(
[
(1, 'Qureshi', 'Imran'),
(2, 'Vidal', 'Michael'),
],
['member_id', 'last_name', 'first_name']
).createOrReplaceTempView("patients")
source_df: DataFrame=spark.table("patients")
df=source_df.select("member_id")
df.createOrReplaceTempView("members")
result_df: DataFrame=mapper.transform(df=df)

Statically Typed Examples

To improve the auto-complete and syntax checking even more, you can define Complex types:

Define a custom data type:

fromspark_auto_mapper.type_definitions.automapper_defined_typesimportAutoMapperTextInputTypefromspark_auto_mapper.helpers.automapper_value_parserimportAutoMapperValueParserfromspark_auto_mapper.data_types.dateimportAutoMapperDateDataTypefromspark_auto_mapper.data_types.listimportAutoMapperListfromspark_auto_mapper_fhir.fhir_types.automapper_fhir_data_type_complex_baseimportAutoMapperFhirDataTypeComplexBaseclassAutoMapperFhirDataTypePatient(AutoMapperFhirDataTypeComplexBase):
# noinspection PyPep8Namingdef__init__(self,
id_: AutoMapperTextInputType,
birthDate: AutoMapperDateDataType,
name: AutoMapperList,
gender: AutoMapperTextInputType
) ->None:
super().__init__()
self.value=dict(
id=AutoMapperValueParser.parse_value(id_),
birthDate=AutoMapperValueParser.parse_value(birthDate),
name=AutoMapperValueParser.parse_value(name),
gender=AutoMapperValueParser.parse_value(gender)
)

Now you get auto-complete and syntax checking:

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapperFhir(
view="members",
source_view="patients",
keys=["member_id"]
).withResource(
resource=F.patient(
id_=A.column("a.member_id"),
birthDate=A.date(
A.column("date_of_birth")
),
name=A.list(
F.human_name(
use="usual",
family=A.column("last_name")
)
),
gender="female"
)
)

Publishing a new package

  1. Edit VERSION to increment the version
  2. Create a new release
  3. The GitHub Action should automatically kick in and publish the package
  4. You can see the status in the Actions tab

About

Fluent API to map data from one view to another in Spark. Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Resources

Contributing

Stars

7 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

Repository files navigation

Build and Test

Upload Python Package

Known Vulnerabilities

SparkAutoMapper

Fluent API to map data from one view to another in Spark.

Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Since this is just Python, you can use any Python editor. Since everything is typed using Python typings, most editors will auto-complete and warn you when you do something wrong

Usage

pip install sparkautomapper

Documentation

https://icanbwell.github.io/SparkAutoMapper/

SparkAutoMapper input and output

You can pass either a dataframe to SparkAutoMapper or specify the name of a Spark view to read from.

You can receive the result as a dataframe or (optionally) pass in the name of a view where you want the result.

Dynamic Typing Examples

Set a column in destination to a text value (read from pass in data frame and return the result in a new dataframe)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to a text value (read from a Spark view and put result in another Spark view)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to an int value

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=1050
)

Copy a column (src1) from source_view to destination view column (dst1)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column("src1")
)

Or you can use the shortcut for specifying a column (wrap column name in [])

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]"
)

Convert data type for a column (or string literal)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
birthDate=A.date(A.column("date_of_birth"))
)

Use a Spark SQL Expression (Any valid Spark SQL expression can be used)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Specify multiple transformations

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]",
birthDate=A.date("[date_of_birth]"),
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Use variables or parameters

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)

Use conditional logic

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)
ifparameters["customer"] =="Microsoft":
mapper=mapper.columns(
important_customer=1,
customer_name=parameters["customer"]
)
returnmapper

Using nested array columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).withColumn(
dst2=A.list(
[
"address1",
"address2"
]
)
)

Using nested struct columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.complex(
use="usual",
family="imran"
)
)

Using lists of structs

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.list(
[
A.complex(
use="usual",
family="imran"
),
A.complex(
use="usual",
family="[last_name]"
)
]
)
)

Executing the AutoMapper

spark.createDataFrame(
[
(1, 'Qureshi', 'Imran'),
(2, 'Vidal', 'Michael'),
],
['member_id', 'last_name', 'first_name']
).createOrReplaceTempView("patients")
source_df: DataFrame=spark.table("patients")
df=source_df.select("member_id")
df.createOrReplaceTempView("members")
result_df: DataFrame=mapper.transform(df=df)

Statically Typed Examples

To improve the auto-complete and syntax checking even more, you can define Complex types:

Define a custom data type:

fromspark_auto_mapper.type_definitions.automapper_defined_typesimportAutoMapperTextInputTypefromspark_auto_mapper.helpers.automapper_value_parserimportAutoMapperValueParserfromspark_auto_mapper.data_types.dateimportAutoMapperDateDataTypefromspark_auto_mapper.data_types.listimportAutoMapperListfromspark_auto_mapper_fhir.fhir_types.automapper_fhir_data_type_complex_baseimportAutoMapperFhirDataTypeComplexBaseclassAutoMapperFhirDataTypePatient(AutoMapperFhirDataTypeComplexBase):
# noinspection PyPep8Namingdef__init__(self,
id_: AutoMapperTextInputType,
birthDate: AutoMapperDateDataType,
name: AutoMapperList,
gender: AutoMapperTextInputType
) ->None:
super().__init__()
self.value=dict(
id=AutoMapperValueParser.parse_value(id_),
birthDate=AutoMapperValueParser.parse_value(birthDate),
name=AutoMapperValueParser.parse_value(name),
gender=AutoMapperValueParser.parse_value(gender)
)

Now you get auto-complete and syntax checking:

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapperFhir(
view="members",
source_view="patients",
keys=["member_id"]
).withResource(
resource=F.patient(
id_=A.column("a.member_id"),
birthDate=A.date(
A.column("date_of_birth")
),
name=A.list(
F.human_name(
use="usual",
family=A.column("last_name")
)
),
gender="female"
)
)

Publishing a new package

  1. Edit VERSION to increment the version
  2. Create a new release
  3. The GitHub Action should automatically kick in and publish the package
  4. You can see the status in the Actions tab

About

Fluent API to map data from one view to another in Spark. Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Resources

Contributing

Stars

7 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages

, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

Repository files navigation

Build and Test

Upload Python Package

Known Vulnerabilities

SparkAutoMapper

Fluent API to map data from one view to another in Spark.

Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Since this is just Python, you can use any Python editor. Since everything is typed using Python typings, most editors will auto-complete and warn you when you do something wrong

Usage

pip install sparkautomapper

Documentation

https://icanbwell.github.io/SparkAutoMapper/

SparkAutoMapper input and output

You can pass either a dataframe to SparkAutoMapper or specify the name of a Spark view to read from.

You can receive the result as a dataframe or (optionally) pass in the name of a view where you want the result.

Dynamic Typing Examples

Set a column in destination to a text value (read from pass in data frame and return the result in a new dataframe)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to a text value (read from a Spark view and put result in another Spark view)

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="hello"
)

Set a column in destination to an int value

Set a column in destination to a text value

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=1050
)

Copy a column (src1) from source_view to destination view column (dst1)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column("src1")
)

Or you can use the shortcut for specifying a column (wrap column name in [])

fromspark_auto_mapper.automappers.automapperimportAutoMappermapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]"
)

Convert data type for a column (or string literal)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
birthDate=A.date(A.column("date_of_birth"))
)

Use a Spark SQL Expression (Any valid Spark SQL expression can be used)

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Specify multiple transformations

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1="[src1]",
birthDate=A.date("[date_of_birth]"),
gender=A.expression(
""" CASE WHEN `Member Sex` = 'F' THEN 'female' WHEN `Member Sex` = 'M' THEN 'male' ELSE 'other' END """
)
)

Use variables or parameters

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)

Use conditional logic

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAdefmapping(parameters: dict):
mapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst1=A.column(parameters["my_column_name"])
)
ifparameters["customer"] =="Microsoft":
mapper=mapper.columns(
important_customer=1,
customer_name=parameters["customer"]
)
returnmapper

Using nested array columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).withColumn(
dst2=A.list(
[
"address1",
"address2"
]
)
)

Using nested struct columns

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.complex(
use="usual",
family="imran"
)
)

Using lists of structs

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapper(
view="members",
source_view="patients",
keys=["member_id"]
).columns(
dst2=A.list(
[
A.complex(
use="usual",
family="imran"
),
A.complex(
use="usual",
family="[last_name]"
)
]
)
)

Executing the AutoMapper

spark.createDataFrame(
[
(1, 'Qureshi', 'Imran'),
(2, 'Vidal', 'Michael'),
],
['member_id', 'last_name', 'first_name']
).createOrReplaceTempView("patients")
source_df: DataFrame=spark.table("patients")
df=source_df.select("member_id")
df.createOrReplaceTempView("members")
result_df: DataFrame=mapper.transform(df=df)

Statically Typed Examples

To improve the auto-complete and syntax checking even more, you can define Complex types:

Define a custom data type:

fromspark_auto_mapper.type_definitions.automapper_defined_typesimportAutoMapperTextInputTypefromspark_auto_mapper.helpers.automapper_value_parserimportAutoMapperValueParserfromspark_auto_mapper.data_types.dateimportAutoMapperDateDataTypefromspark_auto_mapper.data_types.listimportAutoMapperListfromspark_auto_mapper_fhir.fhir_types.automapper_fhir_data_type_complex_baseimportAutoMapperFhirDataTypeComplexBaseclassAutoMapperFhirDataTypePatient(AutoMapperFhirDataTypeComplexBase):
# noinspection PyPep8Namingdef__init__(self,
id_: AutoMapperTextInputType,
birthDate: AutoMapperDateDataType,
name: AutoMapperList,
gender: AutoMapperTextInputType
) ->None:
super().__init__()
self.value=dict(
id=AutoMapperValueParser.parse_value(id_),
birthDate=AutoMapperValueParser.parse_value(birthDate),
name=AutoMapperValueParser.parse_value(name),
gender=AutoMapperValueParser.parse_value(gender)
)

Now you get auto-complete and syntax checking:

fromspark_auto_mapper.automappers.automapperimportAutoMapperfromspark_auto_mapper.helpers.automapper_helpersimportAutoMapperHelpersasAmapper=AutoMapperFhir(
view="members",
source_view="patients",
keys=["member_id"]
).withResource(
resource=F.patient(
id_=A.column("a.member_id"),
birthDate=A.date(
A.column("date_of_birth")
),
name=A.list(
F.human_name(
use="usual",
family=A.column("last_name")
)
),
gender="female"
)
)

Publishing a new package

  1. Edit VERSION to increment the version
  2. Create a new release
  3. The GitHub Action should automatically kick in and publish the package
  4. You can see the status in the Actions tab

About

Fluent API to map data from one view to another in Spark. Uses native Spark functions underneath so it is just as fast as hand writing the transformations.

Resources

Contributing

Stars

7 stars

Watchers

6 watching

Forks

Releases

Packages

Used by

Contributors

Languages