testing-graphframes.ipynb
Depending on how I derive the data frames for edges and vertices I'm getting a completely different result.
When using SQL to derive the data frame the result looks completely wrong.
Running the example seems to work fine:
vertex = spark.createDataFrame([
("a", "Alice", 34),
("b", "Bob", 36),
("c", "Charlie", 30),
("d", "David", 29),
("e", "Esther", 32),
("f", "Fanny", 36),
("g", "Gabby", 60)],
["id", "name", "age"])
vertex.display()
edge = spark.createDataFrame([
("a", "b", "friend"),
("b", "c", "follow"),
("c", "b", "follow"),
("f", "c", "follow"),
("e", "f", "follow"),
("e", "d", "friend"),
("d", "a", "friend"),
("a", "e", "friend")],
["src", "dst", "relationship"])
edge.display()
dbutils.fs.rm("/tmp/graphframes-example-connected-components", True)
sc.setCheckpointDir("/tmp/graphframes-example-connected-components")
graph = GraphFrame(vertex, edge)
result = graph.connectedComponents()
result.select("id", "component").orderBy("component").display()
Result correct:
|id. |component |
|g |146028888064|
|a |412316860416|
|b |412316860416|
|d |412316860416|
|c |412316860416|
|f |412316860416|
|e |412316860416|
However this example with the exact same data doesn't work, the component ID is wrong:
from graphframes import GraphFrame
vertex = spark.sql("""
select
cast(null as string ) as id,
cast(null as string) as name,
cast(null as int) as age
where 1=0 union all
select "a", "Alice" , 34 union all
select "b", "Bob" , 36 union all
select "c", "Charlie" , 30 union all
select "d", "David" , 29 union all
select "e", "Esther" , 32 union all
select "f", "Fanny" , 36 union all
select "g", "Gabby" , 60
""")
vertex.display()
edge = spark.createDataFrame([
("a", "b", "friend"),
("b", "c", "follow"),
("c", "b", "follow"),
("f", "c", "follow"),
("e", "f", "follow"),
("e", "d", "friend"),
("d", "a", "friend"),
("a", "e", "friend")],
["src", "dst", "relationship"])
edge.display()
dbutils.fs.rm("/tmp/graphframes-example-connected-components", True)
sc.setCheckpointDir("/tmp/graphframes-example-connected-components")
graph = GraphFrame(vertex, edge)
result = graph.connectedComponents()
result.select("id", "component").orderBy("component").display()
Result incorrect:
|id |component|
|g |146028888064|
|f |412316860416|
|e |670014898176|
|d |807453851648|
|c |1047972020224|
|b |1382979469312|
|a |1460288880640|
I expected both these results to be identical. The vertex and edge data is exactly the same!
Runtime environment:
Databricks https://docs.databricks.com/aws/en/release-notes/runtime/17.3lts-ml
CPU config not GPU.
testing-graphframes.ipynb
Depending on how I derive the data frames for edges and vertices I'm getting a completely different result.
When using SQL to derive the data frame the result looks completely wrong.
Running the example seems to work fine:
Result correct:
However this example with the exact same data doesn't work, the component ID is wrong:
Result incorrect:
I expected both these results to be identical. The vertex and edge data is exactly the same!
Runtime environment:
Databricks https://docs.databricks.com/aws/en/release-notes/runtime/17.3lts-ml
CPU config not GPU.