Skip to content

bug: Inconsistent Results with connectedComponents #760

Description

@shaunryan

testing-graphframes.ipynb

Depending on how I derive the data frames for edges and vertices I'm getting a completely different result.
When using SQL to derive the data frame the result looks completely wrong.

Running the example seems to work fine:

vertex = spark.createDataFrame([
    ("a", "Alice", 34),
    ("b", "Bob", 36),
    ("c", "Charlie", 30),
    ("d", "David", 29),
    ("e", "Esther", 32),
    ("f", "Fanny", 36),
    ("g", "Gabby", 60)],
    ["id", "name", "age"])
vertex.display()


edge = spark.createDataFrame([
    ("a", "b", "friend"),
    ("b", "c", "follow"),
    ("c", "b", "follow"),
    ("f", "c", "follow"),
    ("e", "f", "follow"),
    ("e", "d", "friend"),
    ("d", "a", "friend"),
    ("a", "e", "friend")], 
    ["src", "dst", "relationship"])
edge.display()


dbutils.fs.rm("/tmp/graphframes-example-connected-components", True)
sc.setCheckpointDir("/tmp/graphframes-example-connected-components")

graph = GraphFrame(vertex, edge)
result = graph.connectedComponents()
result.select("id", "component").orderBy("component").display()

Result correct:

|id.   |component |
|g	|146028888064|
|a	|412316860416|
|b	|412316860416|
|d	|412316860416|
|c	|412316860416|
|f	|412316860416|
|e	|412316860416|

However this example with the exact same data doesn't work, the component ID is wrong:

from graphframes import GraphFrame

vertex = spark.sql("""
select 
  cast(null as string   ) as id,
  cast(null as string) as name,
  cast(null as int) as age
where 1=0 union all

select "a", "Alice"   , 34 union all
select "b", "Bob"     , 36 union all
select "c", "Charlie" , 30 union all
select "d", "David"   , 29  union all
select "e", "Esther"  , 32 union all
select "f", "Fanny"   , 36 union all
select "g", "Gabby"   , 60  

""")
vertex.display()

edge = spark.createDataFrame([
    ("a", "b", "friend"),
    ("b", "c", "follow"),
    ("c", "b", "follow"),
    ("f", "c", "follow"),
    ("e", "f", "follow"),
    ("e", "d", "friend"),
    ("d", "a", "friend"),
    ("a", "e", "friend")], 
    ["src", "dst", "relationship"])
edge.display()


dbutils.fs.rm("/tmp/graphframes-example-connected-components", True)
sc.setCheckpointDir("/tmp/graphframes-example-connected-components")

graph = GraphFrame(vertex, edge)
result = graph.connectedComponents()
result.select("id", "component").orderBy("component").display()

Result incorrect:

|id    |component|
|g	|146028888064|
|f	|412316860416|
|e	|670014898176|
|d	|807453851648|
|c	|1047972020224|
|b	|1382979469312|
|a	|1460288880640|

I expected both these results to be identical. The vertex and edge data is exactly the same!

Runtime environment:

Databricks https://docs.databricks.com/aws/en/release-notes/runtime/17.3lts-ml
CPU config not GPU.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions